Summary
DUET-DINO is a simultaneous cross-view latent world model for robot manipulation introduced by researchers including Nisarga Nilavadi and Rudolf Lioutikov (arXiv:2609.10506). Action-conditioned latent world models enable zero-shot goal-conditioned robot planning, but their predictions of fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. DUET-DINO addresses this by jointly learning action-conditioned predictions from static side-camera and wrist-camera observations through cross-view conditioning, exploiting complementary global scene and gripper-centric information to enable latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, it consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. Trained from scratch on DROID and RoboArena datasets, DUET-DINO generalizes robustly under visual distribution shifts. The authors also show that DINOv3 predictions capture action-conditioned scene changes better than V-JEPA 2, whose wrist-view predictions underestimate fine-grained visual dynamics. Code and checkpoints will be open-sourced.
Paper Overview
Field: Computer Vision (CV)
Authors: Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez, Tobias Jülg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
Published: 2026-09-09
arXiv:
2609.10506Summary
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, the authors introduce
DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning.
By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space.
Key Results
- Consistently outperforms single-view and independent dual-view baselines across:
- Spatially diverse reach tasks: 92% success
- Orientation-intensive angled-reach tasks: 72.5% success
- Multi-goal grasp-and-lift tasks: 60.0% success
- Trained from scratch on DROID and RoboArena datasets
- Generalizes robustly under visual distribution shifts
- Encoder comparison: V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, while DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning.
Resources
- Project page: https://utn-air.github.io/DUET-DINO
- Code and model checkpoints will be open-sourced.
---
*Auto-collected on 2026-09-11*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178634707