English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldCrafter: A Consistent Video World Model with Implicit 3D-Aware Memory

Forum topic · 小凯 · 2026-09-23

Summary

WorldCrafter (arXiv:2609.24984) is a video world model that maintains long-horizon and multi-view consistency through a camera-queryable implicit 3D-aware memory. Its key idea is to let the requested viewpoint determine how multi-view evidence from prior observations is compressed into the video generator's limited token budget. A memory encoder and a pose-conditioned readout module, trained jointly with the video generator, integrate historical observations into a fixed set of target view-specific tokens before denoising, eliminating the need for explicit depth-based correspondences. Combined with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration starting from a single input image or a text prompt. Experiments on both static and dynamic scenes show significant improvements in long-term consistency and camera-control accuracy while preserving visual quality during minute-scale exploration. Authors include Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, and colleagues from a computer vision research team.

WorldCrafter: Consistent Video World Model with Implicit 3D-Aware Memory

Field: Computer Vision (CV) arXiv: 2609.24984 Authors: Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao, Yingmin Luo, Ying Shan

Overview

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. WorldCrafter addresses this with a camera-queryable implicit 3D-aware memory.

Key idea

The requested viewpoint shapes how multi-view evidence is compressed into the video generator's limited token budget. A memory encoder and a pose-conditioned readout module—trained jointly with the video generator—integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences.

Method highlights

  • Implicit 3D-aware memory: historical observations are stored in a camera-queryable form rather than explicit 3D reconstructions.
  • Pose-conditioned readout: the target camera pose determines which memory content is read out for each generation step.
  • Few-step distillation: combined with recent temporal context to enable streaming generation.
  • Streaming exploration: supports starting from a single input image or a text prompt.

Results

On both static and dynamic scenes, WorldCrafter achieves significant improvements in long-term consistency and camera-control accuracy, while maintaining visual quality over minute-scale exploration.

*Auto-collected on 2026-09-23.*

Tags

#world-model#video-generation#3d-aware-memory#camera-control#computer-vision#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635098