Paper Overview
Field: Computer Vision (CV) Authors: Wenxuan Guo, Ziyuan Li, Meng Zhang Release date: 2026-05-25 arXiv: 2505.14489
Abstract (translated)
Vision-Language-Action (VLA) models show strong potential for generalizable robotic manipulation by unifying perception and action. However, existing VLA systems rely primarily on text instructions, making it difficult to resolve spatial ambiguity in complex scenes with multiple similar objects. To address this limitation, we introduce gestures as a parallel instruction modality and propose the Gesture-aware Vision-Language-Action model (GesVLA).
Our method encodes gesture features directly into the latent space so they can participate in both high-level reasoning and low-level action generation, and adopts a dual-VLM architecture to achieve tight coupling between gesture representation and the action policy.
On the data side, we build a scalable gesture data generation pipeline by rendering hand models onto real scene images. This reduces the sim-to-real visual gap while producing rich data with diverse motion patterns and corresponding pointing annotations. We further adopt a two-stage training strategy that equips the model with gesture awareness and action prediction capabilities.
We evaluate the approach on multiple real-world robot tasks, including a controlled block-manipulation task for validation and more practical scenarios such as product and agricultural produce selection. Experimental results show that incorporating gestures consistently improves target grounding accuracy and human-robot interaction efficiency, particularly in complex and cluttered environments.
--- *Auto-collected on 2026-05-25*