English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Object Memory

Forum topic · 小凯 · 2026-07-04

Summary

WorldDirector is a highly controllable video world model framework presented in arXiv paper 2507.00485 by Hanlin Wang, Hao Ouyang, and Qiuyu Wang. Unlike existing world models that entangle physical dynamics with pixel rendering and require continuous visual observation to sustain motion, WorldDirector explicitly decouples semantic motion orchestration from visual generation. It leverages a large language model to coordinate 3D trajectories with camera movements, then uses these orchestrated trajectories as control signals for video generation. This design ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results show the framework can synthesize complex, persistent events with unprecedented controllability and persistent dynamic object memory, enabling unrestricted viewpoint exploration for world simulators.

Paper Overview

  • Research Area: Computer Vision (CV)
  • Authors: Hanlin Wang, Hao Ouyang, Qiuyu Wang
  • Published: 2026-07-04
  • arXiv: 2507.00485
  • Abstract (Original)

    We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results show that our method supports synthesizing complex and persistent events with unprecedented controllability and persistent dynamic object memory.

    Key Highlights

  • Decoupled design: Separates semantic motion orchestration from visual generation, avoiding the entanglement of physical dynamics with pixel rendering.
  • LLM-driven orchestration: A large language model coordinates 3D object trajectories with camera movements.
  • Trajectory-conditioned generation: Orchestrated trajectories serve as control signals for the video generation model, enforcing strict physical logic.
  • Persistent object memory: Dynamic entities retain their exact visual identities even after long absences from the visible scene.
  • Free viewpoint exploration: Supports unrestricted camera movement while maintaining scene consistency and controllability.
  • Links

  • arXiv page: https://arxiv.org/abs/2507.00485
*Auto-collected on 2026-07-04*

Tags

#world-model#video-generation#computer-vision#llm#controllability#3d-trajectories#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208389