Feynman's Letter: Do You Want to Be a "Painting Photocopier" or a "Physics-Savvy Creator"? — On the New Era of Visual Generation
After reading the ten-thousand-word survey Visual Generation in the New Era (arXiv: 2504.19983), I feel we are at the epicenter of a physics revolution called "visual cognition."
To show you why today's AI video generation is no longer just "stitching pixels together," let's talk about the act of "creation."
1. The Status Quo: The "Atomic-Level Mapping" Photocopier
Early visual generation systems (such as early GANs or simple diffusion models) were, at their core, an advanced pixel photocopier.
- The pain point: Give it "a dog running," and it would search its vast training set for pixel blocks related to "dog" and "running" (atomic mapping), then forcibly stitch them together. It doesn't understand gravity or occlusion. This is what's called "visual hallucination lacking physical logic."
- The physical image (world model): Today's top models (like Sora) are no longer pixel movers—they are a sandbox with Newton's laws built in. When it draws a cup falling to the ground, it isn't painting the pixels of "cup" and "shattered glass"; it is simulating the physical process of "gravitational acceleration" and "brittle material collision" in its head.
- The agentic quality: This means visual models are beginning to possess an understanding of the world's causality. They can accept complex interactive instructions and even generate the next frame based on environmental feedback. It is no longer a static canvas—it is an interactive universe.
2. Agentic World Modeling: The Sandbox with a Built-In Physics Engine
This paper reveals the greatest leap happening in visual generation: moving from pixel stitching toward agentic world modeling.
3. The Feynman-Style Judgment: Seeing Is "Reconstructing Physics"
So-called "realism" is not about how high your resolution is.
It is whether the world you render obeys the underlying causal laws that make stars orbit and apples fall.
This survey tells us: the endgame of visual generation is not to replace a few artists, but to reconstruct a physical simulator isomorphic to the real world within a silicon-based one.
When AI can not only "draw" the wind but also "understand" the fluid dynamics of wind, it ceases to be a tool—it becomes our microscope for exploring parallel universes.
Takeaway:
When evaluating any video generation foundation model, don't just look at how beautiful the output is.
Test its "physical common sense."
If a model generates videos where mirrors show no reflections and water ignores gravity, its "realism" is just a probabilistic magic trick; only when its generated pixels strictly obey physical laws does it truly possess the power of "creation."
---
Reference: *Visual Generation in the New Era* (arXiv: 2504.19983)