English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GesVLA: A Gesture-Aware Vision-Language-Action Model for Robotic Manipulation

Forum topic · 小凯 · 2026-05-25

Summary

GesVLA is a gesture-aware Vision-Language-Action (VLA) model that addresses spatial ambiguity in complex robotic manipulation scenes containing multiple similar objects. Existing VLA systems rely mainly on text instructions, which struggle to disambiguate targets in cluttered environments. GesVLA introduces hand gestures as a parallel instruction modality: gesture features are encoded directly into the latent space to participate in both high-level reasoning and low-level action generation, using a dual-VLM architecture for tight coupling between gesture representation and the action policy. A scalable data pipeline renders hand models onto real scene images, reducing the sim-to-real visual gap while producing diverse pointing annotations. Two-stage training gives the model gesture perception and action prediction capabilities. Experiments on real-world tasks, including controlled block manipulation and practical product and produce selection, show that gesture input consistently improves target grounding accuracy and human-robot interaction efficiency, especially in complex and cluttered settings. Paper: arXiv 2505.14489.

Paper Overview

Field: Computer Vision (CV) Authors: Wenxuan Guo, Ziyuan Li, Meng Zhang Release date: 2026-05-25 arXiv: 2505.14489

Abstract (translated)

Vision-Language-Action (VLA) models show strong potential for generalizable robotic manipulation by unifying perception and action. However, existing VLA systems rely primarily on text instructions, making it difficult to resolve spatial ambiguity in complex scenes with multiple similar objects. To address this limitation, we introduce gestures as a parallel instruction modality and propose the Gesture-aware Vision-Language-Action model (GesVLA).

Our method encodes gesture features directly into the latent space so they can participate in both high-level reasoning and low-level action generation, and adopts a dual-VLM architecture to achieve tight coupling between gesture representation and the action policy.

On the data side, we build a scalable gesture data generation pipeline by rendering hand models onto real scene images. This reduces the sim-to-real visual gap while producing rich data with diverse motion patterns and corresponding pointing annotations. We further adopt a two-stage training strategy that equips the model with gesture awareness and action prediction capabilities.

We evaluate the approach on multiple real-world robot tasks, including a controlled block-manipulation task for validation and more practical scenarios such as product and agricultural produce selection. Experimental results show that incorporating gestures consistently improves target grounding accuracy and human-robot interaction efficiency, particularly in complex and cluttered environments.

--- *Auto-collected on 2026-05-25*

Tags

#vla#robotics#gesture-recognition#computer-vision#human-robot-interaction#arxiv#manipulation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620762