论文概要
研究领域: CV
作者: Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez, Tobias Jülg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
发布时间: 2026-09-09
arXiv: 2609.10506
中文摘要
动作条件潜在世界模型预测未来视觉表征,实现零样本目标条件机器人规划与控制。然而,它们对细粒度空间和旋转动作的预测对于完整7自由度末端执行器控制并不可靠。为此,本文提出 DUET-DINO,一种同时跨视角潜在世界模型,通过跨视角条件联合学习静态侧视和腕部相机观测的动作条件预测。利用互补的全局场景和以夹爪为中心的信息,DUET-DINO 实现了完整7自由度动作空间的潜在规划。在空间多样的到达、方向密集的倾斜到达和多目标抓取-提升任务中,DUET-DINO 始终优于单视角和独立双视角基线,到达任务成功率92%、倾斜到达72.5%、提升60.0%。DUET-DINO 在DROID和RoboArena数据集上从头训练,在视觉分布偏移下稳健泛化。
原文摘要
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, DUET-DINO consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. DUET-DINO is trained from scratch on DROID and RoboArena datasets and generalizes robustly under visual distribution shifts. We further show that while V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning. The code and model checkpoints will be open-sourced. Project page: https://utn-air.github.io/DUET-DINO
自动采集于 2026-09-11
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。