English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SANA-WM: An Efficient 2.6B-Parameter Open-Source World Model for Minute-Scale 720p Video Generation

Forum topic · 小凯 · 2026-05-17

Summary

SANA-WM is a 2.6-billion-parameter open-source world model introduced in an arXiv paper (2605.15178) that natively generates high-fidelity, minute-scale (60-second) 720p video with precise camera control. It matches the visual quality of large industrial baselines like LingBot-World and HY-WorldPlay while being far more efficient. Four key designs: (1) hybrid linear attention combining frame-level gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling; (2) dual-branch camera control for accurate 6-DoF trajectory following; (3) a two-stage pipeline with a long-video refiner for cross-sequence consistency; (4) a robust annotation pipeline extracting metric-scale 6-DoF camera poses from public video. Trained in 15 days on 64 H100 GPUs using only ~213K public clips with metric pose supervision, it generates 60-second clips on a single GPU; a distilled NVFP4-quantized variant denoises 60s 720p clips in 34 seconds on a single RTX 5090. It achieves stronger action-following accuracy than prior open baselines and 36x higher throughput in scalable world modeling.

Overview

SANA-WM is an efficient 2.6-billion-parameter open-source world model, natively trained for one-minute generation, synthesizing high-fidelity 720p minute-scale video with precise camera control. It achieves visual quality comparable to large-scale industrial baselines (such as LingBot-World and HY-WorldPlay) while being significantly more efficient.

Paper details

  • Field: Computer Vision (CV)
  • Authors: Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie
  • arXiv: 2605.15178
  • Key Architectural Designs

    Four core designs drive the architecture:

    1. Hybrid linear attention: Combines frame-level gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. 2. Dual-branch camera control: Ensures precise 6-DoF trajectory following. 3. Two-stage generation pipeline: Applies a long-video refiner to first-stage outputs, improving quality and consistency across sequences. 4. Robust annotation pipeline: Extracts precise metric-scale 6-DoF camera poses from public videos to produce high-quality, spatiotemporally consistent action labels.

    Efficiency Highlights

  • Trained with only ~213K public video clips with metric-scale pose supervision.
  • Training completed in 15 days on 64 H100 GPUs.
  • Generates each 60-second clip on a single GPU.
  • The distilled variant can be deployed on a single RTX 5090, denoising a 60-second 720p clip in 34 seconds using NVFP4 quantization.
  • Benchmark Results

    On the paper's one-minute world model benchmark, SANA-WM demonstrates:

  • Stronger action-following accuracy than previous open-source baselines.
  • 36x higher throughput in scalable world modeling while achieving comparable visual quality.
--- *Automatically collected on 2026-05-17.*

Tags

#world-model#video-generation#diffusion-models#linear-attention#camera-control#computer-vision#arxiv#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620166