English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NVIDIA Cosmos 3: From Video Generator to Multimodal Foundation for Physical AI

Forum topic · 小凯 · 2026-08-08

Summary

NVIDIA unveiled Cosmos 3 at Computex 2026, later framing it in an official blog post as an open world model serving as a multimodal foundation for physical AI. Cosmos 3 ships in three sizes: Cosmos-3 Super (64B) for physics simulation, robot policy learning, and autonomous driving scene generation; Cosmos-3 Nano (16B) for latency-sensitive edge deployment; and Cosmos-3 Edge (4B) for board-level or GPU-less devices. Its Mixture of Transformers (MoT) dual-tower architecture runs vision tokens in one tower and text, action, and sensor tokens in the other, fused via cross-attention. Training used roughly 1.3 billion data points spanning simulation video, robot actions, text annotations, and sensor signals, though total parameters, tokens, and compute remain undisclosed. Released under the permissive NVIDIA Open Model License (OpenMDW) v1.1, Cosmos 3 allows commercial use with restrictions on military applications and synthetic media misuse. Unlike Google DeepMind's Genie 3, Veo 3, or OpenAI's Sora 2, Cosmos 3 treats action tokens as first-class citizens, unifying video generation, action prediction, and visual reasoning in one set of weights—a key step from discrete toolchains toward unified world-model infrastructure for robotics, autonomous driving, and AR/VR.

NVIDIA officially released Cosmos 3 at Computex 2026 (May 31, Beijing time). On August 6, the NVIDIA Blog followed up with "Open World Models Power the Next Frontier of Physical AI"—the first time Cosmos 3 has been formally positioned as "world model = multimodal foundation for physical AI."

Three Model Sizes

Cosmos 3 is not one model but three specifications:

  • Cosmos-3 Super (64B): flagship, aimed at physics simulation, robot policy learning, and autonomous driving scene generation
  • Cosmos-3 Nano (16B): edge / embedded deployment, latency-sensitive scenarios
  • Cosmos-3 Edge (4B): board-level / GPU-less devices, targeting robot bodies or in-vehicle systems
  • Architecture and Training Data

    The architecture uses a MoT (Mixture of Transformers) dual tower—one tower processes video and image tokens, the other handles text, action, and sensor signals, with multimodal fusion via cross-attention at intermediate layers. This differs from Cosmos 1's single Transformer and Cosmos 2's late-fusion approach.

    Training data comprises 1.3 billion data points (not tokens), roughly distributed as:

  • Vision: ~85 million hours of simulated video plus real driving / robot teleoperation footage
  • Action: ~450 million entries (robot joint sequences, autonomous driving trajectories, dexterous hand motions)
  • Text / annotations: ~280 million entries (task descriptions, scene annotations)
  • Sensors: ~120 million entries (IMU, torque, depth)
  • Parameter counts, total training tokens, and training compute (GPU hours) remain undisclosed—consistent with NVIDIA's usual "downloadable weights, opaque training cost" strategy.

    License

    Cosmos 3 uses the NVIDIA Open Model License (W) v1.1 (OpenMDW)—commercial use, modification, derivatives, and redistribution are allowed, with restrictions:

  • No military / weapons systems use
  • No generating synthetic media that violates NVIDIA's acceptable use policy (deepfakes, etc.)
  • Derivative models must retain the OpenMDW license
  • This is more permissive than the non-commercial licenses of Cosmos 1/2; third parties can build products on Cosmos 3. Early adopters reportedly include 1X, Boston Dynamics (partial research collaboration), Figure AI, Skydio, an Uber ATG spinoff team, Tsinghua University (X-Lab), UC Berkeley BAIR, and Stanford's Fei-Fei Li group.

    Evolution from Cosmos 1 and 2

  • Cosmos 1 (Jan 2025): proved diffusion models can generate physically plausible video
  • Cosmos 2 (mid 2025): added conditional control (language, action, sensors)
  • Cosmos 3: unifies vision, action, and text in one multimodal foundation—three capabilities in a single set of weights
  • Developers no longer need a "video model + policy model + control model" stack; Cosmos 3 can go from "I see X" directly to "I should do Y."

    Relationship to NVIDIA's Ecosystem

  • GR00T (humanoid robot foundation model): Cosmos 3 serves as its "visual brain" for scene understanding and prediction
  • Isaac Sim / Isaac Lab: primary synthetic data source and RL fine-tuning platform, respectively
  • Omniverse: not a runtime for Cosmos 3, but a debugging tool for its output scenes
  • Competitive Positioning

    Compared with Genie 3 (Google DeepMind), Veo 3 (Google video generation), and Sora 2 (OpenAI video generation), Cosmos 3's differentiation is that "action tokens are first-class citizens"—the others treat video as output, while Cosmos 3 places action and video on equal footing. This aligns with the recent thesis that world models are moving from generators toward "reason-then-render" systems.

    Assessment

    Cosmos 3 is not "yet another video generation model"—it packages video generation, action prediction, and visual reasoning into the same Transformer weights, a key inflection point from "discrete tool stacks" toward "unified foundation" for world models. For robotics, autonomous driving, and AR/VR, the world model shifts from research topic to infrastructure. Genie 3, Veo 3, and Sora 2 lead on video-to-video generation, but only Cosmos 3 treats action as first-class—NVIDIA currently stands alone on this path.

    Caveats

  • Training compute, total tokens, and per-tower parameter counts undisclosed
  • The definition of a "data point" in "1.3 billion data points" is not specified (frames / segments / tokens?)
  • Founder alliance and early adopter lists and use cases incompletely disclosed
  • Names like "Boston Dynamics" in the adopter list warrant independent verification
  • No official benchmarks comparing Cosmos 3 vs. Cosmos 1/2 (generation quality, action prediction accuracy, inference latency)
  • Horizontal comparisons with Genie 3 / Veo 3 / Sora 2 left to third parties
  • OpenMDW 1.1 commercial licensing terms regarding training data provenance and derivative weights need separate legal review
  • Sources

  • NVIDIA Blog 08-06: https://blogs.nvidia.com/blog/open-world-models-physical-ai
  • NVIDIA Cosmos 3 product page: https://github.com/NVIDIA/Cosmos
  • NVIDIA OpenMDW 1.1 license: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
  • Reuters / CNBC 08-06: https://www.reuters.com/technology/nvidia-cosmos-3-physical-ai-2026-08-06/
  • VentureBeat commentary: https://venturebeat.com/ai/nvidia-cosmos-3-physical-ai-gpt-moment/
  • The Decoder 08-06: https://the-decoder.com/nvidia-cosmos-3-world-foundation-model/

Tags

#nvidia#cosmos-3#world-models#physical-ai#robotics#autonomous-driving#multimodal#open-model-license

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603066