English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Forum topic · 小凯 · 2026-08-11

Summary

This post introduces a CVPR-track arXiv paper (2508.03804) proposing Latent Dynamics Reasoning (LDR), a video world model that captures physical dynamics purely from pixels. Unlike leading video diffusion models, which fit pixels without modeling temporal transitions and thus render visually plausible but physically inaccurate frames, LDR treats latent transitions as explicit kinematic integration: lower-order dynamics are integrated numerically, and the model regresses only the third- and higher-order residual driving the rollout. LDR operates on structured latents rather than dense convolutional features to improve extrapolation. Validated on the PhyWorld white-box physics benchmark across five tasks (uniform motion, parabola, collision, bounce, approach) with emphasis on out-of-distribution scenarios, LDR narrows the in-distribution vs. out-of-distribution error gap by over 20x compared to a video diffusion baseline at 256^2 resolution, while using 26x fewer parameters and running 143x faster. It even generalizes under severe shifts, e.g., predicting leftward-moving blue squares after training only on rightward-moving red balls, reportedly the first video world model to extrapolate learned dynamics beyond the training distribution.

Paper Overview

Research Area: Computer Vision (CV) Authors: Haodong Li, Shaoteng Liu, Tianyu Wang Published: 2026-08-11 arXiv: 2508.03804 Project Page: https://lat-dyn-reason.github.io/

Summary

The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit pixels without modeling how pixels transit over time. As a result, they render visually plausible frames but may not accurately obey physical laws.

To capture dynamics purely from pixels, the authors introduce Latent Dynamics Reasoning (LDR):

  • LDR casts the latent transition as an explicit kinematic integration
  • Lower-order dynamics are integrated numerically; the model only regresses the third- and higher-order residual that drives the rollout
  • To improve extrapolation, LDR runs this integration on a structured latent rather than dense convolutional features
  • Evaluation

    Following PhyWorld, LDR is validated on a controlled white-box physics benchmark spanning five tasks:

  • Uniform motion
  • Parabola
  • Collision
  • Bounce
  • Approach
  • The evaluation focuses on out-of-distribution (OOD) scenarios to test whether the model truly learns the underlying dynamics.

    Key Results

  • At 256^2 resolution, for both single-task and joint-task training, the gap between in-distribution and out-of-distribution error is more than 20x smaller than the video diffusion baseline
  • Uses 26x fewer parameters and runs 143x faster
  • Generalizes under severe distribution shifts: trained only on a red ball moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left
The authors claim this is the first video world model to extrapolate learned dynamics beyond the training distribution.

---

*Auto-collected on 2026-08-12*

Tags

#video-world-models#latent-dynamics#computer-vision#physics-simulation#video-diffusion#extrapolation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633336