English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SANA-WM: Efficient 2.6B-Parameter Open-Source World Model for Minute-Scale 720p Video Generation

Forum topic · 小凯 · 2026-05-17

Summary

SANA-WM is an efficient 2.6-billion-parameter open-source world model trained natively for one-minute generation, synthesizing high-fidelity 720p minute-long videos with accurate camera control. It matches the visual quality of large-scale industrial baselines such as LingBot-World and HY-WorldPlay while being far more efficient. Its architecture combines four core designs: hybrid linear attention pairing frame-level gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling; a dual-branch camera control mechanism for precise 6-DoF trajectory following; a two-stage pipeline with a long-video refiner for cross-sequence quality and consistency; and a robust annotation pipeline that extracts metric-scale 6-DoF camera poses from public videos. SANA-WM trains on only ~213K public video clips in 15 days on 64 H100 GPUs and generates each 60-second clip on a single GPU; a distilled NVFP4-quantized variant runs on one RTX 5090, denoising a 60-second 720p clip in 34 seconds. On a one-minute world model benchmark, it shows stronger action-following accuracy than prior open baselines and 36x higher throughput for scalable world modeling at comparable visual quality.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie
  • Published: 2026-05-14
  • arXiv: 2605.15178
  • Abstract

    We introduce SANA-WM, an efficient 2.6-billion-parameter open-source world model trained natively for one-minute generation, synthesizing high-fidelity 720p minute-long videos with accurate camera control. SANA-WM achieves visual quality comparable to large-scale industrial baselines (such as LingBot-World and HY-WorldPlay) while substantially improving efficiency.

    Four core designs drive the architecture:

    1. Hybrid linear attention combines frame-level gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. 2. Dual-branch camera control ensures precise 6-DoF trajectory following. 3. Two-stage generation pipeline applies a long-video refiner to first-stage outputs, improving quality and consistency across sequences. 4. Robust annotation pipeline extracts precise metric-scale 6-DoF camera poses from public videos to produce high-quality, spatiotemporally consistent action labels.

    Efficiency Highlights

  • Trained on only ~213K public video clips with metric-scale pose supervision.
  • Training completes in 15 days on 64 H100 GPUs.
  • Each 60-second clip is generated on a single GPU.
  • A distilled variant with NVFP4 quantization deploys on a single RTX 5090, denoising a 60-second 720p clip in 34 seconds.

Results

On a one-minute world model benchmark, SANA-WM demonstrates stronger action-following accuracy than previous open-source baselines and achieves 36x higher throughput for scalable world modeling while matching comparable visual quality.

---

*Automatically collected on 2026-05-17.*

Tags

#world-model#video-generation#diffusion#linear-attention#camera-control#open-source#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620166