[论文] DepthWorld: 3D World Model for Robot Manipulation
研究领域: CV 作者: Jai Bardhan, Josef Sivic, Vladimir Petrik 发布时间: 2026-10-06 arXiv: 2610.08780
论文概要
研究领域: CV 作者: Jai Bardhan, Josef Sivic, Vladimir Petrik 发布时间: 2026-10-06 arXiv: 2610.08780
中文摘要
世界模型为机器人技术提供了数据驱动的传统仿真器替代方案,应用涵盖策略评估、改进和规划。所有这些用途都依赖于忠实的3D几何,但当前基于视频的世界模型仅在RGB上训练,产生的rollouts逐帧看是正确的,但并不能组成一致的3D世界。弥合这一差距需要在两个方面取得进展:用于操作的大规模3D监督,以及能够在不干扰强预训练视频先验的情况下吸收这些监督的架构。我们引入了一个校准流程,将学习到的立体深度与联合因子图结合,汇集从同一物理机器人收集的所有回合,以恢复其共享的运动学参数以及每场景的外参。应用于DROID数据集,这产生了DROID-3D——一个校准的3D数据集,提供密集的度量深度和重新校准的多视角外参(在外部相机上90%的回合中实现<0.7像素重投影误差)。然后,我们训练DepthWorld,这是一个基于Stable Video Diffusion的世界模型,通过空间潜在平铺联合预测多视角RGB和深度,保持预训练的VAE不变。深度监督在相同训练预算下将RGB预测本身提高了+1.48 dB PSNR(相比相同的仅RGB基线),同时产生准确的度量深度供下游几何推理使用。
原文摘要
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to ...
*自动采集于 2026-10-08*
#论文 #arXiv #CV #小凯