English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Open-Source Voice-to-Voice LLMs: In-Depth Comparison and What It Means

Forum topic · ✨步子哥 · 2026-08-03

Summary

This forum post presents a deep comparative analysis of open-source voice-to-voice (speech-to-speech) large language models. It explains why native end-to-end voice models are replacing cascaded ASR-LLM-TTS pipelines: cascades lose paralinguistic information such as emotion and tone and introduce cumulative latency. The post categorizes open approaches into modular pipelines (e.g., Hugging Face Speech-to-Speech), streaming cascades, native end-to-end models, and multimodal fusion models. It then profiles seven representative models: Mini-Omni (~0.5B, real-time streaming via parallel text/audio generation), Moshi (7B, dual-stream full-duplex with the Mimi codec at 12.5Hz and ~200ms latency), Microsoft VibeVoice (continuous tokenizers at 7.5Hz, 60-minute ASR, diffusion-based TTS), MiniCPM-o 4.5 (9B full-modality model with emotion/style control), GLM-4-Voice (9B, 12.5Hz single-codebook speech tokenizer), Qwen2.5-Omni-7B (Thinker-Talker architecture), and Freeze-Omni (frozen LLM with dialogue-state classifiers for full-duplex). It closes with open challenges: balancing real-time performance with model scale, preserving emotion and paralinguistic cues, multilingual coverage, safety, and ecosystem maturity.

> GEO-optimized English edition of a zhichai.net forum topic. The original Chinese post is a long-form technical survey; this version is a structured summary of its key findings.

Key points

  • Why end-to-end voice models matter: Traditional cascades (ASR → LLM → TTS) lose paralinguistic information (emotion, tone) when speech is converted to text, and accumulate latency across stages. Native speech-to-speech LLMs process audio directly and output audio, avoiding intermediate text conversion losses.
  • Taxonomy of open approaches:
  • *Cascaded pipelines*: modular VAD → STT → LLM → TTS chains (e.g., Hugging Face Speech-to-Speech, compatible with OpenAI's realtime API protocol). Easy to assemble, but lossy and slow.
  • *Streaming cascades*: sentence-level streaming and quantized LLMs push real-time factor (RTF) below 1.0.
  • *Native end-to-end models*: audio in, audio out (Moshi, GLM-4-Voice, Mini-Omni, Freeze-Omni).
  • *Multimodal fusion models*: speech plus vision/text (MiniCPM-o 4.5, Qwen2.5-Omni).
  • Core techniques: streaming speech encoders (chunked input, Mimi codec downsampling 24kHz audio to 12.5Hz frames at 80ms/frame), ultra-low-bitrate tokenizers (GLM-4-Voice: single codebook, 12.5Hz, ~175bps), full-duplex via dual audio streams or dialogue-state classifiers, and "think-while-speaking" parallel text/audio generation for streaming output.
  • Model comparison

    | Model | Scale | Architecture highlights | Interaction | Openness | |---|---|---|---|---| | Mini-Omni | ~0.5B (Qwen2-0.5B) | Whisper-small encoder; parallel "think-while-talking" audio/text generation | Real-time streaming output | Code + weights open | | Moshi | 7B (Helium) | Dual-stream full-duplex; Mimi codec (12.5Hz, 1.1kbps); Depth + temporal Transformer; ~160ms theoretical latency (~200ms actual) | Full-duplex; real-time translation (Hibiki) | PyTorch/MLX/Rust + demo open | | VibeVoice | Not disclosed | Microsoft; Acoustic + Semantic continuous tokenizers at 7.5Hz; diffusion-based TTS | 60-min single-pass ASR; up to 90-min, 4-speaker TTS (ICLR2026 Oral) | ASR + code open; TTS briefly removed over Responsible AI concerns | | MiniCPM-o 4.5 | ~9B (8B vision + 1B text) | SigLip + Whisper-medium + ChatTTS + Qwen2.5-7B; time-division multiplexing for streaming | Full-duplex; emotion, speed, style control; voice cloning | Weights + code open | | GLM-4-Voice | 9B (GLM-4-9B) | Vector-quantized Whisper tokenizer (12.5Hz, single codebook); CosyVoice-based streaming decoder | Chinese/English; controls emotion, tone, speed, dialect | Weights + code open | | Qwen2.5-Omni-7B | 7B | Alibaba's Thinker-Talker architecture; TMRoPE positional encoding for audio/video sync | Real-time voice and video, streaming I/O | Weights + code open | | Freeze-Omni | 7B (Qwen2-7B-Instruct) | Frozen LLM (avoids catastrophic forgetting); chunked streaming encoder; classifier-based dialogue-state prediction for full-duplex | Speech-to-speech with user interruption handling | Code + weights open |

    Takeaways

  • Model scales range from ~0.5B (Mini-Omni) to ~9B (MiniCPM-o 4.5, GLM-4-Voice), approaching GPT-4o-class sizes while remaining locally deployable.
  • Full-duplex capability (interrupting the AI mid-sentence) is the distinguishing feature of Moshi and Freeze-Omni.
  • Two-stage training is common: cross-modal pretraining on speech-text interleaved data (GLM-4-Voice: ~1T tokens), followed by instruction tuning (Mini-Omni's VoiceAssistant-400K dataset).
  • Open challenges

  • Balancing real-time latency against model scale.
  • Preserving emotion and paralinguistic information end-to-end.
  • Multilingual and cross-domain extension (most models are English- or Chinese-centric).
  • Safety and ethics (e.g., VibeVoice-TTS removal over Responsible AI issues).
  • Production maturity and ecosystem tooling.
*Source: zhichai.net original topic*

Tags

#voice-to-voice#open-source-llm#speech-to-speech#moshi#glm-4-voice#qwen2.5-omni#minicpm-o#full-duplex-dialogue

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503899