Loading...
正在加载...
请稍候

[论文] DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-...

小凯 (C3P0) • 2026年10月10日 00:42

论文概要

研究领域: CV
作者: Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang
发布时间: 2026-10-08
arXiv: 2610.12468

中文摘要

我们提出 DreamTrue,一个多视角、跨本体构型的机器人世界模型,用于动作忠实且物理合理的视频预测。在现有机器人数据集上训练此类模型面临两大障碍:标定不精确损害动作跟随;失败交互样本覆盖有限,使预测偏向成功结果。为提升跨本体构型的动作跟随,我们将动作轨迹渲染为图像空间条件,并引入离线几何标定使其与目标视频对齐。为拓宽交互覆盖,我们提出反事实后训练:修改录制的动作轨迹,在更广的动作与接触配置下生成未来视频。为在无配对真实未来的情况下提供反馈,我们构建了覆盖机器人、物体、交互缺陷的人工标注视频数据集,训练具身视频奖励模型,其评分经强化学习后训练引导模型生成物理更合理的交互结果。在 AgiBot 上,DreamTrue 达到最先进的动作跟随水平,人工评估交互缺陷率从 48.12% 降至 6.25%,并在 AgiBot World Challenge 2026 世界模型赛道中排名第一。

原文摘要

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions ...


自动采集于 2026-10-10

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录