English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

Forum topic · 小凯 · 2026-08-28

Summary

WorldDirector is a controllable video world model framework introduced by researchers including Hanlin Wang and Qifeng Chen, presented on arXiv (2607.02517). The framework targets persistent dynamic object memory and unrestricted viewpoint exploration in world simulators. Unlike existing world models that entangle physical dynamics with pixel rendering and depend on continuous visual observation to sustain motion, WorldDirector explicitly decouples semantic motion orchestration from visual generation. It leverages an LLM to coordinate 3D object trajectories with camera movements, then uses these orchestrated trajectories as control signals for video generation. This design enforces strict physical logic and appearance stability, allowing the model to preserve the exact visual identities of dynamic entities even when they re-enter a scene after prolonged absence from view. The work falls under computer vision (cs.CV) and was published on July 2, 2026.

Paper Overview

Research Area: cs.CV

Authors: Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen

Published: 2026-07-02

arXiv: 2607.02517

Abstract

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view.

Key Contributions

  • Persistent dynamic object memory: Dynamic entities retain their visual identities even after being out of view for extended periods.
  • Decoupled architecture: Semantic motion orchestration is separated from visual generation, avoiding the entanglement of physical dynamics with pixel rendering common in prior world models.
  • LLM-driven trajectory orchestration: An LLM coordinates 3D object trajectories together with camera movements.
  • Trajectory-controlled video generation: Orchestrated trajectories serve as control signals for the video generator, enforcing strict physical logic and stable appearance.
  • Unrestricted viewpoint exploration: The simulator supports free camera movement without degrading scene consistency.
  • Links

  • arXiv page: https://arxiv.org/abs/2607.02517
*Auto-collected on 2026-08-28.*

Tags

#world-models#video-generation#llm#computer-vision#controllability#persistent-memory#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634133