English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeCAL: Contact-Aware Latent Co-Imagination for Physically-Grounded Dexterous Vision-Language-Action Models

Forum topic · 小凯 · 2026-09-10

Summary

DeCAL is a physically-grounded dexterous vision-language-action (VLA) model for contact-rich robotic manipulation, introduced in an arXiv paper (2609.09119) by Yankai Fu and colleagues. Dexterous manipulation is challenging for existing VLA models due to severe visual occlusions and complex contact dynamics, and prior tactile approaches rely on homogeneous multimodal fusion without adaptive integration or explicit dynamics modeling. DeCAL is built on a Mixture-of-Transformers (MoT) architecture with specialized experts for understanding, imagination, and action generation. It features Adaptive Visuo-Tactile Fusion, which dynamically regulates tactile interactions via a contact-aware gating strategy, and Visuo-Tactile Latent Co-Imagination, which jointly models visual and tactile dynamics to endow the policy with implicit physical world knowledge. Experiments show state-of-the-art performance with a 71% average success rate and 83.4% progress success rate, plus strong generalization to unseen scenarios.

Overview

Field: cs.RO, cs.AI Authors: Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang Published: 2026-09-08 arXiv: 2609.09119

Abstract (original)

Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation, most approaches still rely on homogeneous multimodal fusion, lacking adaptive tactile integration and explicit modeling of physical dynamics. In this work, we present DeCAL, a physically-grounded dexterous vision-language-action model that unifies understanding, imagination and action generation for contact-rich dexterous manipulation. Built upon a Mixture-of-Transformers (MoT) architecture, DeCAL leverages specialized experts for each capability while enabling efficient information flow among them. To effectively leverage tactile information, we introduce Adaptive Visuo-Tactile Fusion that dynamically regulates tactile interactions via a contact-aware gating strategy. Furthermore, we propose Visuo-Tactile Latent Co-Imagination to jointly model visual and tactile dynamics, equipping the policy with implicit physical world knowledge. Experimental results show that DeCAL consistently achieves state-of-the-art performance across all tasks, attaining a 71% average success rate and an 83.4% progress success rate, while also demonstrating strong generalization to unseen scenarios.

Key Contributions

  • Unified framework: DeCAL unifies understanding, imagination, and action generation for contact-rich dexterous manipulation in a single physically-grounded VLA model.
  • Mixture-of-Transformers architecture: Specialized experts handle each capability while maintaining efficient information flow between them.
  • Adaptive Visuo-Tactile Fusion: A contact-aware gating strategy dynamically regulates tactile interactions, addressing the limitations of homogeneous multimodal fusion.
  • Visuo-Tactile Latent Co-Imagination: Joint modeling of visual and tactile dynamics gives the policy implicit physical world knowledge.
  • Results

  • State-of-the-art performance across all evaluated tasks
  • 71% average success rate
  • 83.4% progress success rate
  • Strong generalization to unseen scenarios

Tags

#robotics#vision-language-action#dexterous-manipulation#tactile-sensing#mixture-of-transformers#arxiv#embodied-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634681