Paper Overview
- Field: Computer Vision (CV)
- Authors: Keming Wu, Zuhao Yang, Kaichen Zhang, Shizun Wang, Haowei Zhu, Sicong Leng, Zhongyu Yang, Qijie Wang, Sudong Wang, Ziting Wang, et al.
- Published: 2026-04-30
- arXiv: 2604.28185
- Flow matching
- Unified understanding-and-generation models
- Improved visual representations
- Post-training and reward modeling
- Data curation and filtering
- Synthetic data distillation
- Sampling acceleration
- Benchmark evaluation
- In-the-wild stress testing
- Expert-constrained case studies
Abstract (English, original)
Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning, persistent state, long-horizon consistency, and causal understanding. We argue that the field should move beyond appearance synthesis toward intelligent visual generation: plausible visuals grounded in structure, dynamics, domain knowledge, and causal relations.
The Five-Level Taxonomy
The paper frames the evolution of visual generation as a progression through five levels, from passive renderers to interactive, agentic, world-aware generators:
1. Atomic Generation 2. Conditional Generation 3. In-Context Generation 4. Agentic Generation 5. World-Modeling Generation
Key Technical Drivers
On Evaluation
The authors caution that current evaluations tend to overestimate progress, placing too much emphasis on perceptual quality while neglecting structural, temporal, and causal defects. To address this, the paper combines:
Takeaway
This roadmap offers a capability-centric perspective for understanding, evaluating, and advancing next-generation intelligent visual generation systems—the shift from appearance synthesis to visuals grounded in structure, dynamics, knowledge, and causality.
---
*Auto-collected on 2026-05-03.*