English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SANA-WM: Efficient 2.6B World Model for Minute-Scale 720p Video Generation with Camera Control

Forum topic · 小凯 · 2026-05-16

Summary

SANA-WM is an efficient 2.6B-parameter open-source world model natively trained for one-minute video generation, producing high-fidelity 720p minute-scale videos with precise camera control. It matches the visual quality of large industrial baselines such as LingBot-World and HY-WorldPlay while being far more efficient. Its architecture combines four key designs: hybrid linear attention pairing frame-wise Gated DeltaNet with softmax attention for memory-efficient long-context modeling, dual-branch camera control for accurate 6-DoF trajectory adherence, a two-stage generation pipeline with a long-video refiner, and a robust annotation pipeline that extracts metric-scale 6-DoF camera poses from public videos. Trained in 15 days on 64 H100 GPUs using roughly 213K annotated public video clips, SANA-WM generates a 60-second clip on a single GPU, and its NVFP4-quantized distilled variant runs on one RTX 5090, denoising a 60-second 720p clip in 34 seconds. It shows stronger action-following accuracy than prior open baselines and comparable visual quality at 36x higher throughput.

Overview

Field: Computer Vision Authors: Haoyi Zhu, Haozhe Liu, Yuyang Zhao arXiv: 2505.08634

SANA-WM is an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. It achieves visual quality comparable to large-scale industrial baselines such as LingBot-World and HY-WorldPlay, while significantly improving efficiency.

Core Design

Four core designs drive the architecture:

1. Hybrid Linear Attention — combines frame-wise Gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. 2. Dual-Branch Camera Control — ensures precise 6-DoF trajectory adherence. 3. Two-Stage Generation Pipeline — applies a long-video refiner to stage-1 outputs, improving quality and consistency across sequences. 4. Robust Annotation Pipeline — extracts accurate metric-scale 6-DoF camera poses from public videos to produce high-quality, spatiotemporally consistent action labels.

Efficiency Highlights

  • Trained on only ~213K public video clips with metric-scale pose supervision.
  • Training completed in 15 days on 64 H100 GPUs.
  • Generates each 60-second clip on a single GPU.
  • Distilled variant with NVFP4 quantization runs on a single RTX 5090, denoising a 60-second 720p clip in 34 seconds.

Results

On the one-minute world model benchmark, SANA-WM demonstrates stronger action-following accuracy than prior open-source baselines and achieves comparable visual quality for scalable world modeling at 36x higher throughput.

Paper: <https://arxiv.org/abs/2505.08634>

Tags

#world-model#video-generation#diffusion#linear-attention#camera-control#efficient-inference#arxiv#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620091