English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

Forum topic · 小凯 · 2026-09-20

Summary

Agile-WAM (arXiv:2609.20761) is a lightweight tactile World Action Model (WAM) for contact-rich robot control, proposed by Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, and Junshan Zhang. Unlike recent tactile WAMs that depend on large pretrained generative backbones, Agile-WAM encodes visual and tactile observations into a shared latent space that serves as the source of a direct vision-tactile-to-action flow-matching process, jointly generating action chunks and future visual/tactile latents. The key insight is that vision and touch operate on different timescales: adjacent visual frames are highly similar, while tactile signals can change abruptly at contact. The authors therefore introduce multi-horizon multi-modal prediction—longer-horizon supervision for visual latents and next-frame prediction for tactile latents—to capture fine-grained contact dynamics. Across 9 simulated tasks and 5 real contact-rich manipulation tasks, Agile-WAM surpasses the strongest baselines with a 29.4% relative improvement in overall real-world success rate and only 11.9 ms inference latency, showing that multi-modal WAMs can be deployed in agile, high-frequency, high-precision robotic control with a compact architecture.

Paper Overview

Field: Machine Learning Authors: Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang Published: 2026-09-17 arXiv: 2609.20761

Summary

World Action Models (WAMs) go beyond conventional visuomotor policies by jointly predicting future world states and robot actions, allowing the policy to learn the physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limits inference efficiency and flexible deployment.

This paper presents Agile-WAM, an agile tactile World Action Model for contact-rich robot control. Agile-WAM encodes visual and tactile observations into a shared latent space, which serves as the source of a direct vision-tactile-to-action flow-matching process that jointly generates latent representations of action chunks and future visual/tactile latents.

Key Insight

Vision and tactile signals operate on fundamentally different timescales: adjacent visual frames tend to be highly similar, whereas tactile signals can change abruptly at the moment of contact.

Multi-Horizon Multi-Modal Prediction

To exploit this observation, the authors introduce multi-horizon multi-modal prediction:

  • Visual latents receive supervision over a longer time horizon.
  • Tactile latents are supervised only with next-frame prediction, capturing fine-grained contact dynamics.
  • Results

    Evaluated on 9 simulated tasks and 5 real-world contact-rich manipulation tasks, Agile-WAM achieves strong performance, surpassing the strongest baselines while maintaining low inference latency:

  • 29.4% relative improvement in overall real-world success rate
  • 11.9 ms inference latency
The results demonstrate that multi-modal WAMs can be realized with lightweight architectures, making them suitable for high-precision, high-frequency robot control.

---

*Auto-collected on 2026-09-20.*

Tags

#robotics#world-action-models#tactile-sensing#flow-matching#machine-learning#contact-rich-manipulation#vision-tactile-fusion#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635002