[论文] DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-...

研究领域: CV 作者: Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang 发布时间: 2026-10-08 arXiv: 2610.12468

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang 发布时间: 2026-10-08 arXiv: 2610.12468

中文摘要

我们提出 DreamTrue,一个多视角、跨本体构型的机器人世界模型,用于动作忠实且物理合理的视频预测。在现有机器人数据集上训练此类模型面临两大障碍:标定不精确损害动作跟随;失败交互样本覆盖有限,使预测偏向成功结果。为提升跨本体构型的动作跟随,我们将动作轨迹渲染为图像空间条件,并引入离线几何标定使其与目标视频对齐。为拓宽交互覆盖,我们提出反事实后训练:修改录制的动作轨迹,在更广的动作与接触配置下生成未来视频。为在无配对真实未来的情况下提供反馈,我们构建了覆盖机器人、物体、交互缺陷的人工标注视频数据集,训练具身视频奖励模型,其评分经强化学习后训练引导模型生成物理更合理的交互结果。在 AgiBot 上,DreamTrue 达到最先进的动作跟随水平,人工评估交互缺陷率从 48.12% 降至 6.25%,并在 AgiBot World Challenge 2026 世界模型赛道中排名第一。

原文摘要

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions ...


*自动采集于 2026-10-10*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens