English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

KineVLA: Kinematics-Aware Vision-Language-Action Models for Fine-Grained Robot Manipulation

Forum topic · 小凯 · 2026-03-19

Summary

KineVLA (arXiv:2503.13845) introduces a novel kinematics-rich vision-language-action (VLA) task in which language commands densely encode kinematic attributes—direction, trajectory, orientation, and relative displacement—at key moments from initiation through completion. Unlike existing action instructions that capture kinematics only coarsely or partially, this formulation supports fine-grained, personalized manipulation where task goals remain fixed while execution trajectories must adapt to instruction-level kinematic specifications. The proposed KineVLA framework explicitly decouples goal-level invariance from kinematics-level variability using a bi-level action representation and bi-level reasoning tokens, serving as explicit intermediate supervision to align language and action. To support the task, the authors built kinematics-aware VLA datasets covering both simulation and real robot platforms with instruction-level kinematic variations and bi-level annotations. Experiments on LIBERO and Realman-75 robots show KineVLA consistently outperforms strong VLA baselines on kinematics-sensitive benchmarks, achieving more precise, controllable, and generalizable manipulation behavior.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Gaoge Han, Zhengqing Gao, Ziwen Li
  • Published: 2025-03-18
  • arXiv: 2503.13845
  • Summary

    This paper introduces a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes—such as direction, trajectory, orientation, and relative displacement—at key moments from initiation through completion. Unlike existing action instructions that capture kinematics only coarsely or partially, this formulation supports fine-grained and personalized manipulation.

    In this setting, task goals remain invariant while execution trajectories must adapt to instruction-level kinematic specifications.

    Proposed Approach

    To address this challenge, the authors propose KineVLA, a vision-language-action framework that:

  • Explicitly decouples goal-level invariance from kinematics-level variability
  • Uses a bi-level action representation and bi-level reasoning tokens as explicit intermediate supervision to align language and action
  • Dataset

    The authors construct kinematics-aware VLA datasets covering both simulation and real robot platforms, featuring:

  • Instruction-level kinematic variations
  • Bi-level annotations

Results

Extensive experiments on LIBERO and Realman-75 robots demonstrate that KineVLA consistently outperforms strong VLA baselines on kinematics-sensitive benchmarks, achieving more precise, controllable, and generalizable manipulation behavior.

---

*Auto-collected on 2026-03-19.*

Tags

#vision-language-action#robotics#kinematics#fine-grained-manipulation#arxiv#computer-vision#libero-benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168901