Summary
Wonder is a general-purpose video world model that enables real-time, camera-controllable world exploration. From a single image or a conditional video, it builds an interactive world users can navigate with camera movements, discovering unseen regions and revisiting previously observed areas over long horizons. The system is a co-design of three components: a novel camera conditioning method based on a dense coordinate field whose renderings provide spatially aligned motion and orientation cues; an efficient sparse-attention memory mechanism for fast, precise retrieval over arbitrarily long generation contexts; and a refined self-forcing-style distillation pipeline that improves the student model's adherence to control signals while preserving the teacher's diverse generation modes and long-term memory. Wonder synthesizes diverse minute-long videos at 16 FPS while maintaining coherent geometry, appearance, and dynamics. Beyond image-to-video generation, it natively supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time. Paper: arXiv 2607.26037.
Paper Overview
Field: Computer Vision (CV)
Authors: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei
arXiv: 2607.26037
Abstract
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon.
Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy:
- Camera conditioning: A novel approach using a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence.
- Sparse-attention memory: An efficient memory mechanism that supports fast and precise memory retrieval over a growing generation context, letting the model selectively attend to a small set of relevant context tokens regardless of the actual context length.
- Distillation refinements: Multiple techniques to fix issues in the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals while preserving the teacher model's diverse generation modes and long-term memory.
Together, these components allow Wonder to synthesize diverse, minute-long videos at
16 FPS while maintaining coherent geometry, appearance, and dynamics over long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, enabling existing dynamic scenes to be re-shot in real time.
---
*Auto-collected on 2026-07-30.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178503793