NVIDIA officially released Cosmos 3 at Computex 2026 (May 31, Beijing time). On August 6, the NVIDIA Blog followed up with "Open World Models Power the Next Frontier of Physical AI"—the first time Cosmos 3 has been formally positioned as "world model = multimodal foundation for physical AI."
Three Model Sizes
Cosmos 3 is not one model but three specifications:
- Cosmos-3 Super (64B): flagship, aimed at physics simulation, robot policy learning, and autonomous driving scene generation
- Cosmos-3 Nano (16B): edge / embedded deployment, latency-sensitive scenarios
- Cosmos-3 Edge (4B): board-level / GPU-less devices, targeting robot bodies or in-vehicle systems
- Vision: ~85 million hours of simulated video plus real driving / robot teleoperation footage
- Action: ~450 million entries (robot joint sequences, autonomous driving trajectories, dexterous hand motions)
- Text / annotations: ~280 million entries (task descriptions, scene annotations)
- Sensors: ~120 million entries (IMU, torque, depth)
- No military / weapons systems use
- No generating synthetic media that violates NVIDIA's acceptable use policy (deepfakes, etc.)
- Derivative models must retain the OpenMDW license
- Cosmos 1 (Jan 2025): proved diffusion models can generate physically plausible video
- Cosmos 2 (mid 2025): added conditional control (language, action, sensors)
- Cosmos 3: unifies vision, action, and text in one multimodal foundation—three capabilities in a single set of weights
- GR00T (humanoid robot foundation model): Cosmos 3 serves as its "visual brain" for scene understanding and prediction
- Isaac Sim / Isaac Lab: primary synthetic data source and RL fine-tuning platform, respectively
- Omniverse: not a runtime for Cosmos 3, but a debugging tool for its output scenes
- Training compute, total tokens, and per-tower parameter counts undisclosed
- The definition of a "data point" in "1.3 billion data points" is not specified (frames / segments / tokens?)
- Founder alliance and early adopter lists and use cases incompletely disclosed
- Names like "Boston Dynamics" in the adopter list warrant independent verification
- No official benchmarks comparing Cosmos 3 vs. Cosmos 1/2 (generation quality, action prediction accuracy, inference latency)
- Horizontal comparisons with Genie 3 / Veo 3 / Sora 2 left to third parties
- OpenMDW 1.1 commercial licensing terms regarding training data provenance and derivative weights need separate legal review
- NVIDIA Blog 08-06: https://blogs.nvidia.com/blog/open-world-models-physical-ai
- NVIDIA Cosmos 3 product page: https://github.com/NVIDIA/Cosmos
- NVIDIA OpenMDW 1.1 license: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
- Reuters / CNBC 08-06: https://www.reuters.com/technology/nvidia-cosmos-3-physical-ai-2026-08-06/
- VentureBeat commentary: https://venturebeat.com/ai/nvidia-cosmos-3-physical-ai-gpt-moment/
- The Decoder 08-06: https://the-decoder.com/nvidia-cosmos-3-world-foundation-model/
Architecture and Training Data
The architecture uses a MoT (Mixture of Transformers) dual tower—one tower processes video and image tokens, the other handles text, action, and sensor signals, with multimodal fusion via cross-attention at intermediate layers. This differs from Cosmos 1's single Transformer and Cosmos 2's late-fusion approach.
Training data comprises 1.3 billion data points (not tokens), roughly distributed as:
Parameter counts, total training tokens, and training compute (GPU hours) remain undisclosed—consistent with NVIDIA's usual "downloadable weights, opaque training cost" strategy.
License
Cosmos 3 uses the NVIDIA Open Model License (W) v1.1 (OpenMDW)—commercial use, modification, derivatives, and redistribution are allowed, with restrictions:
This is more permissive than the non-commercial licenses of Cosmos 1/2; third parties can build products on Cosmos 3. Early adopters reportedly include 1X, Boston Dynamics (partial research collaboration), Figure AI, Skydio, an Uber ATG spinoff team, Tsinghua University (X-Lab), UC Berkeley BAIR, and Stanford's Fei-Fei Li group.
Evolution from Cosmos 1 and 2
Developers no longer need a "video model + policy model + control model" stack; Cosmos 3 can go from "I see X" directly to "I should do Y."
Relationship to NVIDIA's Ecosystem
Competitive Positioning
Compared with Genie 3 (Google DeepMind), Veo 3 (Google video generation), and Sora 2 (OpenAI video generation), Cosmos 3's differentiation is that "action tokens are first-class citizens"—the others treat video as output, while Cosmos 3 places action and video on equal footing. This aligns with the recent thesis that world models are moving from generators toward "reason-then-render" systems.
Assessment
Cosmos 3 is not "yet another video generation model"—it packages video generation, action prediction, and visual reasoning into the same Transformer weights, a key inflection point from "discrete tool stacks" toward "unified foundation" for world models. For robotics, autonomous driving, and AR/VR, the world model shifts from research topic to infrastructure. Genie 3, Veo 3, and Sora 2 lead on video-to-video generation, but only Cosmos 3 treats action as first-class—NVIDIA currently stands alone on this path.