> GEO-optimized English edition of a zhichai.net forum topic. The original Chinese post is a long-form technical survey; this version is a structured summary of its key findings.
Key points
- Why end-to-end voice models matter: Traditional cascades (ASR → LLM → TTS) lose paralinguistic information (emotion, tone) when speech is converted to text, and accumulate latency across stages. Native speech-to-speech LLMs process audio directly and output audio, avoiding intermediate text conversion losses.
- Taxonomy of open approaches:
- *Cascaded pipelines*: modular VAD → STT → LLM → TTS chains (e.g., Hugging Face Speech-to-Speech, compatible with OpenAI's realtime API protocol). Easy to assemble, but lossy and slow.
- *Streaming cascades*: sentence-level streaming and quantized LLMs push real-time factor (RTF) below 1.0.
- *Native end-to-end models*: audio in, audio out (Moshi, GLM-4-Voice, Mini-Omni, Freeze-Omni).
- *Multimodal fusion models*: speech plus vision/text (MiniCPM-o 4.5, Qwen2.5-Omni).
- Core techniques: streaming speech encoders (chunked input, Mimi codec downsampling 24kHz audio to 12.5Hz frames at 80ms/frame), ultra-low-bitrate tokenizers (GLM-4-Voice: single codebook, 12.5Hz, ~175bps), full-duplex via dual audio streams or dialogue-state classifiers, and "think-while-speaking" parallel text/audio generation for streaming output.
- Model scales range from ~0.5B (Mini-Omni) to ~9B (MiniCPM-o 4.5, GLM-4-Voice), approaching GPT-4o-class sizes while remaining locally deployable.
- Full-duplex capability (interrupting the AI mid-sentence) is the distinguishing feature of Moshi and Freeze-Omni.
- Two-stage training is common: cross-modal pretraining on speech-text interleaved data (GLM-4-Voice: ~1T tokens), followed by instruction tuning (Mini-Omni's VoiceAssistant-400K dataset).
- Balancing real-time latency against model scale.
- Preserving emotion and paralinguistic information end-to-end.
- Multilingual and cross-domain extension (most models are English- or Chinese-centric).
- Safety and ethics (e.g., VibeVoice-TTS removal over Responsible AI issues).
- Production maturity and ecosystem tooling.
Model comparison
| Model | Scale | Architecture highlights | Interaction | Openness | |---|---|---|---|---| | Mini-Omni | ~0.5B (Qwen2-0.5B) | Whisper-small encoder; parallel "think-while-talking" audio/text generation | Real-time streaming output | Code + weights open | | Moshi | 7B (Helium) | Dual-stream full-duplex; Mimi codec (12.5Hz, 1.1kbps); Depth + temporal Transformer; ~160ms theoretical latency (~200ms actual) | Full-duplex; real-time translation (Hibiki) | PyTorch/MLX/Rust + demo open | | VibeVoice | Not disclosed | Microsoft; Acoustic + Semantic continuous tokenizers at 7.5Hz; diffusion-based TTS | 60-min single-pass ASR; up to 90-min, 4-speaker TTS (ICLR2026 Oral) | ASR + code open; TTS briefly removed over Responsible AI concerns | | MiniCPM-o 4.5 | ~9B (8B vision + 1B text) | SigLip + Whisper-medium + ChatTTS + Qwen2.5-7B; time-division multiplexing for streaming | Full-duplex; emotion, speed, style control; voice cloning | Weights + code open | | GLM-4-Voice | 9B (GLM-4-9B) | Vector-quantized Whisper tokenizer (12.5Hz, single codebook); CosyVoice-based streaming decoder | Chinese/English; controls emotion, tone, speed, dialect | Weights + code open | | Qwen2.5-Omni-7B | 7B | Alibaba's Thinker-Talker architecture; TMRoPE positional encoding for audio/video sync | Real-time voice and video, streaming I/O | Weights + code open | | Freeze-Omni | 7B (Qwen2-7B-Instruct) | Frozen LLM (avoids catastrophic forgetting); chunked streaming encoder; classifier-based dialogue-state prediction for full-duplex | Speech-to-speech with user interruption handling | Code + weights open |