ByteDance Seed Releases SeedRealtime, a Native Audio-Visual Full-Duplex Model
Most "real-time voice" models remain half-duplex: they depend on an external VAD to decide turns, alternating question and answer. ByteDance's Seed team's new release moves "when to speak" inside the model itself.
On August 5, ByteDance Seed released SeedRealtime, a natively audio-visual full-duplex large model. It fuses audio, video, and text into a single architecture and interacts in real time over a continuous multimodal stream — "watching, listening, and speaking" simultaneously.
Technical approach
- End-to-end unified modeling, not cascading. The official announcement explicitly calls out the latency accumulation and information loss of ASR+VLM+TTS pipelines, and notes that many end-to-end models still rely on external VAD — making them effectively half-duplex.
- Internalized turn-taking. SeedRealtime treats turn decisions as its own multimodal decision-making, natively supporting real-time interruption and follow-up.
- Streaming architecture. Continuous audio-video chunks in, streaming generation out, with efficiency from quantization and inference optimizations.
- No technical report, no arXiv paper, no open weights, no API (no Volcano Engine integration announcement). Third parties cannot reproduce or independently evaluate.
- "Halved" comes from the vendor's own human evaluation; the evaluation set size and comparison configuration are undisclosed.
- Circulating figures such as "37% fluency gain," "92% airport recognition," and "millisecond-level latency" have no source on the official pages. Content-farm and AI-generated articles repeating them are untrustworthy.
- Most importantly: the official announcement never mentions robots, embodied intelligence, or on-body deployment. The current form factor is video calling inside a phone app only — it should not be described as "embodied robotics deployed."
- Official Chinese project page: https://seed.bytedance.com/zh/SeedRealtime
- Official English blog: https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction
- Official English project page: https://seed.bytedance.com/en/SeedRealtime
Availability and lineage
SeedRealtime is fully rolled out in the Doubao app ("phone call" in the chat box → video call entry). ByteDance claims it is "the first in the industry to achieve large-scale deployment of audio-visual full-duplex technology." Its predecessor is Seeduplex (released April 9), a full-duplex speech model with a claimed +12% fluency improvement, also deployed in Doubao; SeedRealtime extends it from pure speech to audio-video.
What is (and is not) quantified
The only official quantitative metric: in end-to-end human evaluation, conversational pacing problems were halved versus cascaded models, with notably fewer interruptions by the model, lag, and false triggers. There are no latency figures in milliseconds, no parameter counts, and no public benchmark scores.
Seven real-scenario demos accompany the launch: recognizing people and voices at a four-person dinner, ordering in English at a Sichuan restaurant, proactive reminders during a museum tour, real-time correction while operating a coffee machine, co-reading a ResNet paper down to Section 3.4, noise-robust operation at Beijing Daxing Airport without false triggers, and children's English tutoring.
Embodied-AI speculation is inference, not official claims
Continuous visual-stream understanding + proactive reminders + timing judgment + tool calling are precisely the "interaction layer" needed for embodied agents. The official outlook says "from being able to converse to being able to act," and Seed already has a GR-RL VLA robotics line that appears complementary on the capability stack — but this release does not connect or link the two lines in any way. That linkage is the editor's inference, not an official claim.