Paper Overview
Field: Computer Vision (CV) Authors: Yuyang Yin, Zixiang Li, Longxuan Deng Published: 2026-08-13 arXiv: 2508.03421
Motivation
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design, allowing creators to iteratively refine scenes, actions, camera work, and spatiotemporal dynamics. However, existing generative approaches rely on simple prompts and one-shot image or video synthesis to jointly control all of these factors, resulting in weak controllability and limited support for iterative editing.
The authors' key observation: the world consists of multiple elements with geometry, appearance, and other attributes, plus a camera. Different frames arise from local modifications or recompositions of this shared state, with the rest largely reused. What is missing, they argue, is an explicit and persistent working state.
The StateFlow Framework
StateFlow is a state-centric generative previsualization framework. Instead of one-shot video generation, it organizes scene structure, evolution, and cameras using an editable 3D world, while leveraging off-the-shelf video models to enhance visual fidelity when needed. The world is maintained as a persistent, structured 3D state of scene elements and camera configurations, serving as the core working representation.
Three Stages
1. State Construction — Lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization. 2. State Evolution — Translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration on every edit. 3. State Access — Refines camera plans into visually feasible trajectories using rendered feedback reflection, rather than relying solely on VLM semantics.
Results
Experiments demonstrate that StateFlow can generate high-quality 3D worlds suitable for video creation and game-like prototyping.
--- *Auto-collected on 2026-08-14.*