Paper Information
Original title: Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning Authors: Haodong Li, Shaoteng Liu, Tianyu Wang, et al. Institutions: UCSD, Adobe arXiv: 2608.09926
---
An Everyday Analogy
Imagine a baby watching a ball roll. At first, the baby just passively receives pixel changes—the red ball moves from here to there. But gradually, the baby begins to understand: the ball rolls because of an initial push, it decelerates due to friction, and it bounces when it hits a wall.
This is the leap from a "video generator" to a "world model"—not memorizing how pixels change, but understanding why they change.
LDR (Latent Dynamics Reasoning) aims to achieve exactly this leap.
---
Step-by-Step Breakdown
Stage 1: The Limitations of Video Diffusion Models
Current video diffusion models (such as Sora and Runway) are excellent "pixel painters." They learn:
> "Given these frames, what does the next frame probably look like?"
But this learning is statistical, not causal. The model may generate visually plausible frames whose object motion violates physics:
- Balls accelerating with no force applied
- Momentum not conserved after collisions
- Objects appearing or vanishing from nowhere
- Centroid: where the object's center is
- Extent: how large the object is
- acceleration ← previous acceleration + high-order residual
- velocity ← previous velocity + acceleration
- position ← previous position + velocity 3. Learns high-order residuals: the only part requiring a neural network—third-order and higher dynamic changes
- Trained only on a red ball moving left → right
- Correctly predicts: a blue square moving right → left, and even Pikachu's motion
Stage 2: Structured Latent Space
LDR's key insight: don't model dynamics directly at the pixel level; first compress each image into a structured geometric representation.
Concretely, a CNN extracts feature maps, and an "edge soft-argmax" extracts, per channel:
This is like a minimal "physics-engine representation"—keeping only position and size, discarding irrelevant information such as texture, color, and semantics.
Stage 3: Explicit Kinematic Integration
Given three initial frames, LDR does not directly predict the fourth. Instead it:
1. Initializes: computes velocity and acceleration via finite differences 2. Runs an integration chain: explicitly performs kinematic updates
This is essentially teaching the model an approximation of Newton's second law, rather than having it memorize trajectories.
---
The Science
Core Formulas
Structured latent:
Kinematic integration (per step):
where \(f_\theta\) is a lightweight MLP that learns only the high-order residual.
PhyWorld Benchmark
Validated on five physics tasks (uniform motion, parabolic motion, collision, bouncing, approach):
| Metric | LDR | Video diffusion baseline | |--------|-----|--------------------------| | ID–OOD error gap | 20x smaller | Large | | Parameter count | 26x fewer | Many | | Inference speed | 143x faster | Slow |
Generalization
---
Closing Thoughts
> "LDR is like handing AI the key to a physics lab. It no longer just watches movies of the world—it starts to understand why the world works the way it does. When a model that has only ever seen a red ball rolling rightward can predict a blue square flying leftward, that is not memorization; that is comprehension. If Newton were alive, he might smile."
*Review by: Xiaokai | Feynman-style deep dive*
#papers #arXiv #CV #world-models #physics-reasoning #xiaokai