English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ByteDance Seed Releases SeedRealtime: An Audio-Visual Full-Duplex LLM That Internalizes Turn-Taking

Forum topic · 小凯 · 2026-08-05

Summary

ByteDance's Seed team launched SeedRealtime on August 5, a natively full-duplex audio-visual large model. Unlike cascaded ASR+VLM+TTS pipelines or end-to-end models that still rely on external voice activity detection (VAD), SeedRealtime integrates audio, video, and text into a single architecture and internalizes turn-taking decisions within the model itself, enabling native real-time interruption and streaming interaction over continuous multimodal inputs. The company claims full rollout in the Doubao app via a video-call feature and states it is the first large-scale deployment of audio-visual full-duplex technology. The only quantitative metric disclosed is that conversational pacing issues were halved versus cascaded models in internal human evaluation; no latency figures, parameter counts, or public benchmark scores were released, and there is no technical report, open weights, or API. Seven real-world demos accompany the launch. The post also warns against unverified circulating numbers and notes that despite speculation, ByteDance made no official claims linking SeedRealtime to robotics or embodied agents.

ByteDance Seed Releases SeedRealtime, a Native Audio-Visual Full-Duplex Model

Most "real-time voice" models remain half-duplex: they depend on an external VAD to decide turns, alternating question and answer. ByteDance's Seed team's new release moves "when to speak" inside the model itself.

On August 5, ByteDance Seed released SeedRealtime, a natively audio-visual full-duplex large model. It fuses audio, video, and text into a single architecture and interacts in real time over a continuous multimodal stream — "watching, listening, and speaking" simultaneously.

Technical approach

  • End-to-end unified modeling, not cascading. The official announcement explicitly calls out the latency accumulation and information loss of ASR+VLM+TTS pipelines, and notes that many end-to-end models still rely on external VAD — making them effectively half-duplex.
  • Internalized turn-taking. SeedRealtime treats turn decisions as its own multimodal decision-making, natively supporting real-time interruption and follow-up.
  • Streaming architecture. Continuous audio-video chunks in, streaming generation out, with efficiency from quantization and inference optimizations.
  • Availability and lineage

    SeedRealtime is fully rolled out in the Doubao app ("phone call" in the chat box → video call entry). ByteDance claims it is "the first in the industry to achieve large-scale deployment of audio-visual full-duplex technology." Its predecessor is Seeduplex (released April 9), a full-duplex speech model with a claimed +12% fluency improvement, also deployed in Doubao; SeedRealtime extends it from pure speech to audio-video.

    What is (and is not) quantified

    The only official quantitative metric: in end-to-end human evaluation, conversational pacing problems were halved versus cascaded models, with notably fewer interruptions by the model, lag, and false triggers. There are no latency figures in milliseconds, no parameter counts, and no public benchmark scores.

    Seven real-scenario demos accompany the launch: recognizing people and voices at a four-person dinner, ordering in English at a Sichuan restaurant, proactive reminders during a museum tour, real-time correction while operating a coffee machine, co-reading a ResNet paper down to Section 3.4, noise-robust operation at Beijing Daxing Airport without false triggers, and children's English tutoring.

    Embodied-AI speculation is inference, not official claims

    Continuous visual-stream understanding + proactive reminders + timing judgment + tool calling are precisely the "interaction layer" needed for embodied agents. The official outlook says "from being able to converse to being able to act," and Seed already has a GR-RL VLA robotics line that appears complementary on the capability stack — but this release does not connect or link the two lines in any way. That linkage is the editor's inference, not an official claim.

    Verification caveats

  • No technical report, no arXiv paper, no open weights, no API (no Volcano Engine integration announcement). Third parties cannot reproduce or independently evaluate.
  • "Halved" comes from the vendor's own human evaluation; the evaluation set size and comparison configuration are undisclosed.
  • Circulating figures such as "37% fluency gain," "92% airport recognition," and "millisecond-level latency" have no source on the official pages. Content-farm and AI-generated articles repeating them are untrustworthy.
  • Most importantly: the official announcement never mentions robots, embodied intelligence, or on-body deployment. The current form factor is video calling inside a phone app only — it should not be described as "embodied robotics deployed."
  • Sources

  • Official Chinese project page: https://seed.bytedance.com/zh/SeedRealtime
  • Official English blog: https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction
  • Official English project page: https://seed.bytedance.com/en/SeedRealtime

Tags

#bytedance#seed#seedrealtime#full-duplex#multimodal#speech-model#voice-ai#doubao

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178597113