[论文] DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulati...

研究领域: CV 作者: Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan 发布时间: 2026-09-21 arXiv: 2609.24976

论文概要

研究领域: CV 作者: Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan 发布时间: 2026-09-21 arXiv: 2609.24976

中文摘要

灵巧操作依赖于接触动力学,而这些动力学往往无法仅通过视觉完全观测。最近的世界-动作模型(WAM)将预测性视频世界建模与动作生成相结合,但仍以视觉为中心,无法直接建模接触动力学。我们提出 DexTacWAM,一种视触觉 WAM,独立编码每个指尖的触觉信号,通过手指和姿态感知的触觉压缩器聚合特征,并将触觉隐变量注入视频扩散世界模型中,实现联合视触觉世界建模。在 22 自由度双臂平台上的六项接触密集型灵巧操作任务中,DexTacWAM 在每项任务上都取得了最高分,平均得分 70.6,而最强基线仅为 38.0。消融实验表明,性能提升源于将接触演化建模为预测世界状态的一部分,而非仅作为触觉条件:去除触觉世界建模后,四项任务均值从 74.7 降至 26.6,尽管保留了相同的触觉特征和动作专家。通过在冻结的预训练视觉 VAE 上进行四小时的触觉编码器适配,我们的持续视觉到触觉学习方法将预训练视频模型扩展到触觉领域,每个任务仅需约 100 个演示且无需触觉预训练,同时视觉预测质量保持在纯视觉版本 0.5 dB 以内。压缩器保留了 89.4% 的融合前接触召回率,同时实现了 2.26 倍更快的训练和 1.29 倍更快的推理。这些结果表明,预训练视频先验可以以数据和计算高效的方式扩展到分布式多指接触动力学。

原文摘要

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to mod...


*自动采集于 2026-09-23*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens