Summary
GesVLA (arXiv:2505.17381) is a gesture-aware vision-language-action (VLA) model that addresses spatial ambiguity in robot manipulation when multiple similar objects are present. Existing VLA systems rely mainly on text instructions, which struggle to disambiguate targets in cluttered scenes. GesVLA introduces hand gestures as a parallel instruction modality: gesture features are encoded directly into the latent space so they contribute to both high-level reasoning and low-level action generation, using a dual-VLM architecture that tightly couples gesture representations with the action policy. The authors also build a scalable gesture data pipeline by rendering hand models onto real-world images, reducing the sim-to-real visual gap while producing diverse motion patterns with pointing annotations. A two-stage training strategy equips the model with both gesture perception and action prediction. Experiments on real-world tasks, including controlled block manipulation and practical product and produce selection scenarios, show that gesture integration consistently improves object grounding accuracy and human-robot interaction efficiency, especially in complex and cluttered environments.
Paper Overview
- Field: Computer Vision / Robotics
- Authors: Wenxuan Guo, Ziyuan Li, Meng Zhang
- Published: 2025-05-23
- arXiv: 2505.17381
Summary
Vision-language-action (VLA) models show strong potential for general robot manipulation by unifying perception and action. However, existing VLA systems rely primarily on text instructions, which makes it difficult to resolve spatial ambiguity in complex scenes where multiple similar objects coexist.
To address this limitation, the authors introduce gestures as a parallel instruction modality and propose GesVLA (Gesture-aware Vision-Language-Action model). Key contributions:
- Latent-space gesture encoding: Gesture features are embedded directly into the latent space, enabling participation in both high-level reasoning and low-level action generation.
- Dual-VLM architecture: A two-VLM design tightly couples gesture representations with the action policy.
- Scalable gesture data pipeline: Hand models are rendered onto real-world scene images, reducing the sim-to-real visual gap while yielding rich data with diverse motion patterns and corresponding pointing annotations.
- Two-stage training: The model acquires both gesture perception and action prediction capability.
Evaluation
The method was evaluated on multiple real-world robot tasks, including:
- A controlled block-manipulation task for validation
- More practical product selection and agricultural produce selection scenarios
Results show that incorporating gestures consistently improves
object grounding accuracy and
human-robot interaction efficiency, particularly in complex and cluttered environments.
---
*Originally posted on zhichai.net, auto-collected 2026-05-23.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620661