English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Wonder: A Real-Time Camera-Controllable Video World Model

Forum topic · 小凯 · 2026-07-30

Summary

Wonder is a general-purpose video world model that enables real-time, camera-controllable world exploration. From a single image or a conditional video, it builds an interactive world users can navigate with camera movements, discovering unseen regions and revisiting previously observed areas over long horizons. The system is a co-design of three components: a novel camera conditioning method based on a dense coordinate field whose renderings provide spatially aligned motion and orientation cues; an efficient sparse-attention memory mechanism for fast, precise retrieval over arbitrarily long generation contexts; and a refined self-forcing-style distillation pipeline that improves the student model's adherence to control signals while preserving the teacher's diverse generation modes and long-term memory. Wonder synthesizes diverse minute-long videos at 16 FPS while maintaining coherent geometry, appearance, and dynamics. Beyond image-to-video generation, it natively supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time. Paper: arXiv 2607.26037.

Paper Overview

Field: Computer Vision (CV) Authors: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei arXiv: 2607.26037

Abstract

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon.

Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy:

  • Camera conditioning: A novel approach using a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence.
  • Sparse-attention memory: An efficient memory mechanism that supports fast and precise memory retrieval over a growing generation context, letting the model selectively attend to a small set of relevant context tokens regardless of the actual context length.
  • Distillation refinements: Multiple techniques to fix issues in the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals while preserving the teacher model's diverse generation modes and long-term memory.
Together, these components allow Wonder to synthesize diverse, minute-long videos at 16 FPS while maintaining coherent geometry, appearance, and dynamics over long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, enabling existing dynamic scenes to be re-shot in real time.

---

*Auto-collected on 2026-07-30.*

Tags

#video-world-model#computer-vision#camera-controllable-generation#sparse-attention#distillation#interactive-world-simulation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503793