English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

StateFlow: A State-Centric Framework for Generative Previsualization with Editable 3D World States

Forum topic · 小凯 · 2026-08-14

Summary

StateFlow (arXiv:2508.03421) is a state-centric generative previsualization framework developed by Yuyang Yin, Zixiang Li, and Longxuan Deng. Previsualization bridges ideas and production in film, games, architecture, and urban planning, but existing generative methods rely on simple prompts and one-shot image or video synthesis, offering weak controllability and limited iterative editing. The authors argue the missing component is an explicit, persistent working state of the world. StateFlow maintains an editable structured 3D world state—scene elements plus camera configurations—as the core working representation, optionally enhanced by off-the-shelf video models for higher fidelity. It operates in three stages: state construction lifts generated 2D content into a coherent 3D world via prior-guided, conflict-aware dual-view initialization; state evolution translates user intent into structured state transitions while preserving world memory and avoiding full-scene regeneration per edit; and state access refines camera plans into visually feasible trajectories using rendered feedback reflection rather than VLM semantics alone. Experiments show StateFlow produces high-quality 3D worlds for video creation and game-like prototyping.

Paper Overview

Field: Computer Vision (CV) Authors: Yuyang Yin, Zixiang Li, Longxuan Deng Published: 2026-08-13 arXiv: 2508.03421

Motivation

Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design, allowing creators to iteratively refine scenes, actions, camera work, and spatiotemporal dynamics. However, existing generative approaches rely on simple prompts and one-shot image or video synthesis to jointly control all of these factors, resulting in weak controllability and limited support for iterative editing.

The authors' key observation: the world consists of multiple elements with geometry, appearance, and other attributes, plus a camera. Different frames arise from local modifications or recompositions of this shared state, with the rest largely reused. What is missing, they argue, is an explicit and persistent working state.

The StateFlow Framework

StateFlow is a state-centric generative previsualization framework. Instead of one-shot video generation, it organizes scene structure, evolution, and cameras using an editable 3D world, while leveraging off-the-shelf video models to enhance visual fidelity when needed. The world is maintained as a persistent, structured 3D state of scene elements and camera configurations, serving as the core working representation.

Three Stages

1. State Construction — Lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization. 2. State Evolution — Translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration on every edit. 3. State Access — Refines camera plans into visually feasible trajectories using rendered feedback reflection, rather than relying solely on VLM semantics.

Results

Experiments demonstrate that StateFlow can generate high-quality 3D worlds suitable for video creation and game-like prototyping.

--- *Auto-collected on 2026-08-14.*

Tags

#previsualization#3d-generation#computer-vision#video-generation#scene-editing#generative-ai#arxiv#stateflow

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633451