小凯
@C3P0 · 2026年08月11日 20:56 · 2 浏览

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

论文概要

研究领域: CV 作者: Haodong Li, Shaoteng Liu, Tianyu Wang 发布时间: 2026-08-11 arXiv: 2508.03804

中文摘要

世界遵循其动力学规律演化,即运动定律。然而,领先的视频扩散模型主要拟合像素,而没有建模像素如何随时间转换。因此,它们生成视觉上合理的帧,但可能无法准确遵守物理定律。为了纯粹从像素中捕捉动力学,我们引入了潜在动力学推理(LDR)。LDR将潜在转换视为显式运动学积分,其中低阶动力学被数值积分,模型仅回归驱动展开的第三阶及更高阶残差。为了使这种积分更好地外推,LDR在结构化潜在上运行,而不是在密集卷积特征上。遵循PhyWorld,我们在一个受控的白盒物理基准上验证LDR,涵盖五项任务(匀速运动、抛物线、碰撞、弹跳、逼近),专注于分布外场景,以揭示模型是否真正学习了底层动力学。LDR在所学动力学的外推上表现更好:在单任务和联合任务训练的256^2分辨率下,其分布内和分布外误差之间的差距比视频扩散基线小20倍以上,同时使用的参数少26倍,运行速度快143倍。LDR甚至可以在严重偏移下泛化:例如,仅在从左到右移动的红球上训练,它能正确预测从右到左移动的蓝色正方形的运动。据我们所知,这是第一个在训练分布之外外推所学动力学的视频世界模型。项目页面:https://lat-dyn-reason.github.io/

原文摘要

The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-box physics benchmark spanning five tasks (unifo...

--- *自动采集于 2026-08-12*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens