Overview
SANA-WM is an efficient 2.6-billion-parameter open-source world model, natively trained for one-minute generation, synthesizing high-fidelity 720p minute-scale video with precise camera control. It achieves visual quality comparable to large-scale industrial baselines (such as LingBot-World and HY-WorldPlay) while being significantly more efficient.
Paper details
- Field: Computer Vision (CV)
- Authors: Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie
- arXiv: 2605.15178
- Trained with only ~213K public video clips with metric-scale pose supervision.
- Training completed in 15 days on 64 H100 GPUs.
- Generates each 60-second clip on a single GPU.
- The distilled variant can be deployed on a single RTX 5090, denoising a 60-second 720p clip in 34 seconds using NVFP4 quantization.
- Stronger action-following accuracy than previous open-source baselines.
- 36x higher throughput in scalable world modeling while achieving comparable visual quality.
Key Architectural Designs
Four core designs drive the architecture:
1. Hybrid linear attention: Combines frame-level gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. 2. Dual-branch camera control: Ensures precise 6-DoF trajectory following. 3. Two-stage generation pipeline: Applies a long-video refiner to first-stage outputs, improving quality and consistency across sequences. 4. Robust annotation pipeline: Extracts precise metric-scale 6-DoF camera poses from public videos to produce high-quality, spatiotemporally consistent action labels.
Efficiency Highlights
Benchmark Results
On the paper's one-minute world model benchmark, SANA-WM demonstrates: