English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Forum topic · 小凯 · 2026-08-11

Summary

A new paper (arXiv:2508.03804) by Haodong Li, Shaoteng Liu, and Tianyu Wang introduces Latent Dynamics Reasoning (LDR), an approach that captures physical dynamics purely from pixels. Unlike leading video diffusion models, which fit pixels without modeling temporal transitions and thus produce visually plausible but physically inaccurate frames, LDR treats latent transitions as explicit kinematic integration: lower-order dynamics are numerically integrated, and the model regresses only the third- and higher-order residual driving the rollout. LDR operates on structured latents rather than dense convolutional features to improve extrapolation. Evaluated on the PhyWorld white-box physics benchmark across five tasks (uniform motion, projectile, collision, bouncing, and approach) with an emphasis on out-of-distribution scenarios, LDR shows over 20x smaller in-distribution vs. out-of-distribution error gaps than video diffusion baselines at 256^2 resolution, using 26x fewer parameters and running 143x faster. It even generalizes under severe shifts, predicting motion of unseen objects from training on single-object scenarios, arguably making it the first video world model to extrapolate learned dynamics beyond the training distribution. Project page: https://lat-dyn-reason.github.io/

Paper Overview

Field: Computer Vision Authors: Haodong Li, Shaoteng Liu, Tianyu Wang Published: 2026-08-11 arXiv: 2508.03804

Abstract

The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws.

To capture the dynamics purely from pixels, the authors introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features.

Following PhyWorld, LDR is validated on a controlled white-box physics benchmark spanning five tasks (uniform motion, projectile, collision, bouncing, and approach), focusing on out-of-distribution scenarios to reveal whether the model truly learns the underlying dynamics.

Key Results

  • LDR performs better at extrapolating learned dynamics: at 256^2 resolution under both single-task and joint-task training, the gap between its in-distribution and out-of-distribution errors is more than 20x smaller than that of video diffusion baselines.
  • It uses 26x fewer parameters and runs 143x faster than the baselines.
  • LDR generalizes under severe shifts: trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left.
  • To the authors' knowledge, this is the first video world model to extrapolate learned dynamics beyond the training distribution.
  • Links

  • arXiv: https://arxiv.org/abs/2508.03804
  • Project page: https://lat-dyn-reason.github.io/
--- *Auto-collected on 2026-08-12*

Tags

#world-models#video-generation#latent-dynamics#computer-vision#physics-simulation#extrapolation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633358