English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GesVLA: Gesture-Aware Vision-Language-Action Model with Embedded Gesture Representations

Forum topic · 小凯 · 2026-05-25

Summary

GesVLA introduces gesture as a parallel instruction modality to address spatial ambiguity in Vision-Language-Action (VLA) models for general-purpose robotic manipulation. Conventional VLA systems rely on text instructions and struggle in scenes containing multiple similar objects. The authors embed gesture features directly into the latent space, enabling them to participate in both high-level reasoning and low-level action generation via a dual-VLM architecture that tightly couples gesture representation with the action policy. To produce training data, they render parametric hand models into real scene images, reducing the sim-to-real visual gap while generating diverse motion patterns and pointing annotations. A two-stage training strategy teaches the model to interpret gestures and predict actions. Experiments span controlled block-manipulation tasks and practical scenarios such as product and produce selection, showing consistent gains in target-localization accuracy and human-robot interaction efficiency, especially in cluttered and complex environments.

GesVLA: Gesture-Aware Vision-Language-Action Model with Embedded Gesture Representations

  • Research Area: Computer Vision (CV)
  • Authors: Wenxuan Guo, Ziyuan Li, Meng Zhang
  • Released: 2026-05-25
  • arXiv: 2505.14489
  • Summary

    Vision-Language-Action (VLA) models unify perception and action and have shown strong potential for general-purpose robotic manipulation. However, existing VLA systems depend mainly on text instructions and have difficulty resolving spatial ambiguity in complex scenes with multiple similar objects. To overcome this limitation, the authors introduce gesture as a parallel instruction modality and propose the Gesture-Aware Vision-Language-Action model (GesVLA).

    Key Contributions

  • Gesture in the latent space: Gesture features are encoded directly into the model's latent space so they can participate in both high-level reasoning and low-level action generation.
  • Dual-VLM architecture: A tightly coupled design connects gesture representation with the action policy.
  • Scalable gesture data pipeline: Parametric hand models are rendered into real scene images, reducing the sim-to-real visual gap while producing diverse motion patterns and corresponding pointing annotations.
  • Two-stage training strategy: The model first learns gesture awareness, then learns action prediction conditioned on it.
  • Experiments

    The approach is evaluated on multiple real-world robotic tasks:

  • Controlled block-manipulation tasks used for validation.
  • Practical scenarios such as product selection and produce selection.
  • Results

    Incorporating gestures consistently improves:

  • Target localization accuracy.
  • Human-robot interaction efficiency.
The gains are most pronounced in complex and cluttered environments where text-only instructions are ambiguous.

--- *Auto-collected 2026-05-25*

Tags

#robotics#vision-language-action#gesture-recognition#manipulation#dual-vlm#sim-to-real#human-robot-interaction#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620762