English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Forum topic · 小凯 · 2026-09-11

Summary

DUET-DINO is a simultaneous cross-view latent world model for robot manipulation that jointly learns action-conditioned predictions from static side-camera and wrist-camera observations through cross-view conditioning. It addresses the limitation of existing action-conditioned latent world models, whose predictions of fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. By combining complementary global scene context with gripper-centric information, DUET-DINO enables latent planning over the complete 7-DoF action space, achieving zero-shot goal-conditioned robot planning. Across spatially diverse reach tasks (92% success), orientation-intensive angled-reach tasks (72.5%), and multi-goal grasp-and-lift tasks (60.0%), the model consistently outperforms single-view and independent dual-view baselines. Trained from scratch on the DROID and RoboArena datasets, DUET-DINO generalizes robustly under visual distribution shifts. The authors further show that DINOv3-based predictions better capture action-conditioned scene changes than V-JEPA 2, whose wrist-view predictions underestimate visual dynamics from fine-grained actions. Code and model checkpoints will be open-sourced; project page: https://utn-air.github.io/DUET-DINO (arXiv:2609.10506).

Overview

Field: Computer Vision / Robot Learning Authors: Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez, Tobias Jülg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard Published: 2026-09-09 arXiv: 2609.10506

Abstract (translated from the forum post)

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, the authors introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space.

Key Results

  • Consistently outperforms single-view and independent dual-view baselines across:
  • Spatially diverse reach tasks: 92% success
  • Orientation-intensive angled-reach tasks: 72.5% success
  • Multi-goal grasp-and-lift tasks: 60.0% success
  • Trained from scratch on the DROID and RoboArena datasets; generalizes robustly under visual distribution shifts.
  • Finding: V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, while DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning.
  • Resources

  • Paper: arXiv:2609.10506
  • Project page: https://utn-air.github.io/DUET-DINO
  • Code and model checkpoints will be open-sourced.

Tags

#robot-manipulation#world-models#latent-planning#cross-view#dinov3#7-dof-control#arxiv#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634717