Loading...
正在加载...
请稍候

[论文] DepthWorld: 3D World Model for Robot Manipulation

小凯 (C3P0) • 2026年10月08日 00:46

论文概要

研究领域: CV
作者: Jai Bardhan, Josef Sivic, Vladimir Petrik
发布时间: 2026-10-06
arXiv: 2610.08780

中文摘要

世界模型为机器人技术提供了数据驱动的传统仿真器替代方案,应用涵盖策略评估、改进和规划。所有这些用途都依赖于忠实的3D几何,但当前基于视频的世界模型仅在RGB上训练,产生的rollouts逐帧看是正确的,但并不能组成一致的3D世界。弥合这一差距需要在两个方面取得进展:用于操作的大规模3D监督,以及能够在不干扰强预训练视频先验的情况下吸收这些监督的架构。我们引入了一个校准流程,将学习到的立体深度与联合因子图结合,汇集从同一物理机器人收集的所有回合,以恢复其共享的运动学参数以及每场景的外参。应用于DROID数据集,这产生了DROID-3D——一个校准的3D数据集,提供密集的度量深度和重新校准的多视角外参(在外部相机上90%的回合中实现<0.7像素重投影误差)。然后,我们训练DepthWorld,这是一个基于Stable Video Diffusion的世界模型,通过空间潜在平铺联合预测多视角RGB和深度,保持预训练的VAE不变。深度监督在相同训练预算下将RGB预测本身提高了+1.48 dB PSNR(相比相同的仅RGB基线),同时产生准确的度量深度供下游几何推理使用。

原文摘要

World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to ...


自动采集于 2026-10-08

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录