[论文] DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-...
研究领域: CV 作者: Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang 发布时间: 2026-10-08 arXiv: 2610.12468
论文概要
研究领域: CV 作者: Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang 发布时间: 2026-10-08 arXiv: 2610.12468
中文摘要
我们提出 DreamTrue,一个多视角、跨本体构型的机器人世界模型,用于动作忠实且物理合理的视频预测。在现有机器人数据集上训练此类模型面临两大障碍:标定不精确损害动作跟随;失败交互样本覆盖有限,使预测偏向成功结果。为提升跨本体构型的动作跟随,我们将动作轨迹渲染为图像空间条件,并引入离线几何标定使其与目标视频对齐。为拓宽交互覆盖,我们提出反事实后训练:修改录制的动作轨迹,在更广的动作与接触配置下生成未来视频。为在无配对真实未来的情况下提供反馈,我们构建了覆盖机器人、物体、交互缺陷的人工标注视频数据集,训练具身视频奖励模型,其评分经强化学习后训练引导模型生成物理更合理的交互结果。在 AgiBot 上,DreamTrue 达到最先进的动作跟随水平,人工评估交互缺陷率从 48.12% 降至 6.25%,并在 AgiBot World Challenge 2026 世界模型赛道中排名第一。
原文摘要
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions ...
*自动采集于 2026-10-10*
#论文 #arXiv #CV #小凯