Paper Overview
Field: Computer Vision Authors: Haodong Li, Shaoteng Liu, Tianyu Wang Published: 2026-08-11 arXiv: 2508.03804
Abstract
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws.
To capture the dynamics purely from pixels, the authors introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features.
Following PhyWorld, LDR is validated on a controlled white-box physics benchmark spanning five tasks (uniform motion, projectile, collision, bouncing, and approach), focusing on out-of-distribution scenarios to reveal whether the model truly learns the underlying dynamics.
Key Results
- LDR performs better at extrapolating learned dynamics: at 256^2 resolution under both single-task and joint-task training, the gap between its in-distribution and out-of-distribution errors is more than 20x smaller than that of video diffusion baselines.
- It uses 26x fewer parameters and runs 143x faster than the baselines.
- LDR generalizes under severe shifts: trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left.
- To the authors' knowledge, this is the first video world model to extrapolate learned dynamics beyond the training distribution.
- arXiv: https://arxiv.org/abs/2508.03804
- Project page: https://lat-dyn-reason.github.io/