English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

TempoVLA: Speed-Controllable Vision-Language-Action Policies for Robot Manipulation

Forum topic · 小凯 · 2026-06-07

Summary

TempoVLA (arXiv:2606.06491) is a vision-language-action (VLA) model that lets robots explicitly control their execution speed during manipulation tasks. Robot manipulation alternates between low-risk transit phases requiring fast motion and high-risk contact phases requiring slow, precise movement, but existing VLAs inherit only a single fixed speed from training demonstrations. TempoVLA observes that the magnitude of predicted actions already governs robot speed, and turns this into an explicit speed-conditioning mechanism. It combines two components: data-side Variabl-e Speed Trajectory Augmentation (VSTA), which retimes demonstrations to arbitrary target speeds via merging or splitting actions while preserving motion semantics, and a model-side conditioning mechanism that feeds the desired speed into the policy. Experiments in simulation and the real world show bidirectional flexible speed control, improved default 1x performance through better data utilization, and, when paired with a large multimodal model, dynamic speed control that accelerates in low-risk phases and decelerates in high-risk contact phases.

Overview

Field: Machine Learning / Robotics Authors: Dong Jing, Jingchen Nie, Tianqi Zhang Published: 2026-06-04 arXiv: 2606.06491

Summary

Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and deceleration remains almost unexplored.

The authors observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. They turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition.

TempoVLA combines two coupled components:

1. Data-side Variable Speed Trajectory Augmentation (VSTA): retimes demonstrations to arbitrary target speeds by merging or splitting actions, while preserving motion semantics. Statistics show VSTA reaches target speeds with negligible motion error. 2. Model-side conditioning mechanism: feeds the desired speed into the policy.

Results

  • Simulation and real-world experiments show TempoVLA achieves bidirectional, flexible speed control.
  • VSTA also improves default 1x-speed performance through better data utilization.
  • Paired with a large multimodal model, TempoVLA enables dynamic speed control: accelerating during low-risk transit phases and decelerating during high-risk contact phases.
*Auto-collected on 2026-06-07.*

Tags

#tempovla#vision-language-action#robot-manipulation#machine-learning#speed-control#trajectory-augmentation#robotics#vla-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980915