English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NVIDIA NemotronLabs VoiceChat 11B: The First Open-Source Full-Duplex Voice Agent Foundation Model with Tool Calling

Forum topic · 小凯 · 2026-08-10

Summary

NVIDIA released NemotronLabs VoiceChat 11B on Hugging Face, an open full-duplex speech-to-speech foundation model aimed at voice agent developers rather than end users. It replaces the traditional cascaded ASR + LLM + TTS stack with a unified network: a Fast Conformer streaming encoder, a Nemotron Nano v2 9B LLM backbone, and an NVIDIA TTS decoder, achieving a measured smooth turn-taking latency of 448 ms. It is reportedly the first open-source full-duplex model to support tool calling, emitting <TOOLCALL> blocks on a dedicated output channel with configurable on-hold filler speech to keep conversations natural while tools execute. Benchmarks include Full-Duplex-Bench 1.0 (smooth turn-taking TOR 0.82 / 448 ms) and BFCL-v3 voice averaging 56.1%. Deployment requires a single 80 GB GPU (A100-H200/B-series/RTX 6000), x86_64 Linux, and vLLM. The OpenMDW-1.1 license is research-only; commercial licensing is separate. Known limits: 2-minute audio context, max 5 tools per session, no parallel tool calls, no user interruption during tool execution, and ASCII-only system prompts — a significant drawback for Chinese-language use given its English-centric training (~550k hours of audio).

NVIDIA placed NemotronLabs VoiceChat 11B on Hugging Face on August 9. Its target audience is not end consumers — it is a foundation model for teams building voice agents. On the full-duplex front, open-source solutions have long been stuck cycling through three states: "can listen but not speak," "can speak but not listen," or "can do both but cannot call tools." VoiceChat 11B fills that third gap: it can invoke external tools mid-conversation, push <TOOLCALL> blocks out over a bypass channel, and have the agent deliver "on-hold" filler speech during the seconds a tool waits for results, so the conversation never goes cold.

How does it get latency down to 448 ms? By cutting the entire cascaded stack — ASR, LLM, TTS — and replacing it with one unified network: a Fast Conformer encoder (from Nemotron-Speech-Streaming-En-0.6b, continuously encoding a 16 kHz audio stream) + a Nemotron Nano v2 9B LLM backbone (consuming audio tokens, predicting text tokens) + NVIDIA's own TTS decoder and codec (outputting 22.05 kHz audio). No token conversions or API handshakes in between; end-to-end smooth turn-taking latency measured at 448 ms.

Several engineering details are worth unpacking:

The first open-source full-duplex model with tool calling — previous open-source efforts scored zero on this front. VoiceChat 11B opens a dedicated output channel where tool call scripts are emitted as <TOOLCALL>[{"name":..., "arguments":...}]</TOOLCALL>, and business code returns results via <TOOL_RESPONSE>[...]</TOOL_RESPONSE>. The key design is "on-hold speech": each tool can have a pre-configured line spoken at invocation time, so users don't assume the call dropped while an API runs for 2 seconds. This lets the conversation "naturally pretend" to wait for async results — a genuine step forward in conversational system usability.

Architecture components — four pieces assembled: Fast Conformer encoder + Nemotron Nano v2 backbone + TTS decoder + a dedicated tool-calling output channel. Training data is roughly 550,000 hours of audio (real + synthetic), building on research from SALM-Duplex and Audio Flamingo 3.

Benchmarks — Full-Duplex-Bench 1.0: smooth turn-taking TOR 0.82 / 448 ms; user-interruption TOR 1.00 / 480 ms; pause-handling TOR (lower is better) synthetic 0.153 / Candor 0.255. AU Harness BFCL-v3 voice: simple 58.5%, multiple 62.5%, parallel 42.5%, parallel-multiple 27.5%, irrelevance 89.6%, average 56.1%. Full-Duplex-Bench v3 tool calling: tool selection 82.5%, argument accuracy 44.2%, pass@1 33%. NVIDIA self-reports rank 2 among open full-duplex models on VoiceBench and rank 2 on Full-Duplex-Bench 1.0 open-source.

Deployment requirements and limitations — the boundaries really worth highlighting:

  • The license is OpenMDW-1.1, officially stated as "research purposes only," with commercial licensing to be discussed separately. NVIDIA offers no hosted API itself; NGC provides a container and HF provides weights, and that's it.
  • A single 80 GB GPU (A100 / H100 / H200 / B100 / B200 / RTX 6000), x86_64 Linux, vLLM runtime — teams without GPU resources cannot touch this today.
  • Context cap of 2 minutes of audio; degrades into unrecoverable gibberish after several turns; occasional runaway self-talk after a turn ends; occasional dropped words in user speech transcription.
  • Hard tool-calling limits: max 5 tools per session; no parallel calls; users cannot interrupt during tool execution (to preserve causal closure of <TOOLCALL> / <TOOL_RESPONSE>).
  • System prompts and tool responses must be ASCII-only and TTS-friendly — developers must write prompts without emoji, full-width characters, or Unicode quotes. This is a clear shortcoming for Chinese scenarios: VoiceChat 11B's Fast Conformer is an English streaming model, and the training set is predominantly English.
  • Three judgments to put on the table first.

    First, this is a "foundation model," not a "product." With 448 ms latency, an 80 GB single GPU, and a research-only license, it is not for consumers today — it is a reference and baseline for engineering teams stacking voice agents. It pushes open-source full-duplex from "demonstrable" to "runnable in evals," a direct stimulus to researchers working on Moshi, Audio Flamingo 3, SALM-Duplex, and similar competitors.

    Second, on-hold speech is an underrated design point. Traditional cascaded stacks can only "freeze and wait" during tool calls, breaking the conversational experience. VoiceChat 11B fills that conversational gap with a preset line — acknowledging a fact: the usability bottleneck of voice agents is not latency itself, but the moment when "the user thinks the other side hung up." This follows the same paradigm shift in the Agent era of speech/multimodal systems as Cosmos 3 making "action tokens" first-class citizens (8-08 topicId 178603066) and SeedRealtime internalizing "when to speak" into the model (8-06 topicId 178597113) — beyond latency, agents need explicit rules for "in-conversation behavior."

    Third, the shape of the full-duplex track in H2 2026 is getting clearer. On the open-source side, NVIDIA has completed the tool-calling line with NemotronLabs; on the commercial side, GPT-4o Realtime, Gemini Live, and Realtime API are pushing hosted versions. Open-source foundations with tool calling + reliable commercial managed delivery will coexist. For domestic (Chinese) teams, NVIDIA has put two transferable pain points on the table: the boundaries of OpenMDW commercial licensing, and the ASCII-only system prompt restriction on Chinese interactions.

    Who it suits: voice AI teams with 80 GB GPU R&D capacity (in-car voice, customer service CX, game NPCs, retail kiosks, accessibility). Who it doesn't: consumer products needing immediate production, businesses needing deep Chinese support, and rapid integrators needing turnkey hosted APIs.

    Sources:

  • Official model card (Hugging Face): https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
  • Official code repo (NeMo Speech branch): https://github.com/NVIDIA-NeMo/Speech/tree/nemotron-labs-voicechat
  • Official NGC container: https://catalog.ngc.nvidia.com/orgs/nim/nvidia/containers/nemotron-labs-voicechat
  • OpenMDW 1.1 license: https://github.com/OpenMDW/OpenMDW/blob/main/1.1/LICENSE.OpenMDW-1.1
  • MarkTechPost technical breakdown: https://www.marktechpost.com/2026/08/09/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech-model-with-450-ms-turn-taking-and-live-tool-calling
  • Full-Duplex-Bench 1.0 paper: https://arxiv.org/abs/2503.04721
  • AU Harness BFCL-v3: https://github.com/ServiceNow/AU-Harness
  • VoiceBench paper: https://arxiv.org/abs/2410.17196

Tags

#nvidia#voice-agents#full-duplex#speech-to-speech#open-source-models#tool-calling#llm#speech-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178630988