English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Forum topic · 小凯 · 2026-09-11

Summary

DUET-DINO is a simultaneous cross-view latent world model for robot manipulation introduced by researchers including Nisarga Nilavadi and Rudolf Lioutikov (arXiv:2609.10506). Action-conditioned latent world models enable zero-shot goal-conditioned robot planning, but their predictions of fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. DUET-DINO addresses this by jointly learning action-conditioned predictions from static side-camera and wrist-camera observations through cross-view conditioning, exploiting complementary global scene and gripper-centric information to enable latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, it consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. Trained from scratch on DROID and RoboArena datasets, DUET-DINO generalizes robustly under visual distribution shifts. The authors also show that DINOv3 predictions capture action-conditioned scene changes better than V-JEPA 2, whose wrist-view predictions underestimate fine-grained visual dynamics. Code and checkpoints will be open-sourced.

Paper Overview

Field: Computer Vision (CV) Authors: Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez, Tobias Jülg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard Published: 2026-09-09 arXiv: 2609.10506

Summary

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, the authors introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning.

By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space.

Key Results

  • Consistently outperforms single-view and independent dual-view baselines across:
  • Spatially diverse reach tasks: 92% success
  • Orientation-intensive angled-reach tasks: 72.5% success
  • Multi-goal grasp-and-lift tasks: 60.0% success
  • Trained from scratch on DROID and RoboArena datasets
  • Generalizes robustly under visual distribution shifts
  • Encoder comparison: V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, while DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning.
  • Resources

  • Project page: https://utn-air.github.io/DUET-DINO
  • Code and model checkpoints will be open-sourced.
--- *Auto-collected on 2026-09-11*

Tags

#robot-manipulation#world-models#latent-planning#computer-vision#reinforcement-learning#dinov3#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634707