Loading...
正在加载...
请稍候

[论文] MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction

小凯 (C3P0) 2026年07月21日 00:44

论文概要

研究领域: cs.CV
作者: Homanga Bharadhwaj, Yash Jangir
发布时间: 2026-07-21
arXiv: 2507.15491

中文摘要

人类能够通过被动观察推断物体可能的运动方式:杯子可能被拿起,抽屉可能滑动,盖子可能旋转关闭。这类预测揭示了在现实世界中行动所需的物理交互后果。我们研究如何从普通的人物交互单目视频中学习这种预判能力。给定一段简短的视频上下文,MotionForesight预测被操作物体上各点的未来3D轨迹。这将交互预测转化为以物体为中心的3D运动预测,且不对物体属性做任何假设。我们的核心洞见是:视频预测模型已经编码了关于人类交互过程中物体运动的丰富先验知识。我们将这些先验从像素预测重新导向未来的3D场景流。我们从基于预训练视频模型的密集3D追踪器出发,从完整片段生成伪真值轨迹,并仅使用观察到的帧来训练预测器。我们用学习的掩码潜变量替代未来的RGB和几何信息,并训练一个轻量级适配器将回顾性追踪表示转化为前向预测器,同时冻结大型视频和追踪组件。仅使用4万个人类视频且无需语言等辅助输入,MotionForesight就能泛化到各种分布外的物体、环境、视角和交互。它还显著优于使用超过百万训练视频的更大模型。这些结果表明,我们可以高效地将视频先验重新用于具身智能的显式几何预测。

原文摘要

Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical consequences of interaction needed to act in the real world. We study how to learn this anticipation from ordinary monocular videos of human-object interaction. Given a short observed video context, MotionForesight predicts future 3D trajectories for points on the manipulated object. This casts interaction prediction as object-centered 3D motion forecasting without any assumptions on the object properties. Our key insight is that video prediction models already encode rich priors about how objects move during human interactions. We redirect these priors from pixel prediction toward future 3D scene flow. We start from a dense 3D tracker built on a pretrained video model, generate pseudo-ground-truth tracks from complete clips, and train the forecaster using only the observed frames. We replace future RGB and geometry with learned mask latents and train a lightweight adapter to turn the retrospective tracking representation into a forward predictor, while freezing the large video and tracking components. Using just 40k human videos and no auxiliary inputs such as language, MotionForesight generalizes across diverse out-of-distribution objects, environments, viewpoints, and interactions. It also outperforms substantially larger models that use over a million training videos. These results show that we can efficiently re-purpose video priors into explicit geometric forecasts for embodied intelligence.


自动采集于 2026-07-21

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录