English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LaST-R1: Reinforcing Action via Adaptive Physical Latent Reasoning for Vision-Language-Action Models

Forum topic · 小凯 · 2026-05-03

Summary

LaST-R1 is a unified Vision-Language-Action (VLA) framework for robotic manipulation that integrates latent Chain-of-Thought (CoT) reasoning over physical dynamics before action execution, paired with a dedicated reinforcement learning post-training paradigm. The authors introduce Latent-to-Action Policy Optimization (LAPO), a novel RL algorithm that jointly optimizes the latent reasoning process and action generation, improving physical-world modeling and robustness in interactive environments. An adaptive latent CoT mechanism allows the policy to dynamically adjust reasoning depth based on task complexity. Unlike prior approaches limited to static imitation learning or pure action-space RL, LaST-R1 addresses latency and discretization issues of explicit linguistic reasoning while optimizing the underlying physical reasoning process. Experiments show LaST-R1 achieves a near-perfect 99.8% average success rate on the LIBERO benchmark with only one supervised warm-up stage, surpassing state-of-the-art methods in both convergence speed and performance. Real-world deployment demonstrates up to 44% improvement over the initial policy across four complex tasks, including single-arm and dual-arm setups. Paper: arXiv:2604.28192.

Paper Overview

Field: Computer Vision / Robotics Authors: Hao Chen, Jiaming Liu, Zhonghao Yan, Nuowei Han, Renrui Zhang, Chenyang Gu, Jialin Gao, Ziyu Guo, Siyuan Qian, Yinxi Wang, Peng Jia, Chi-Wing Fu, Shanghang Zhang, Pheng-Ann Heng Published: 2026-04-30 arXiv: 2604.28192

Summary

Vision-Language-Action (VLA) models have increasingly incorporated reasoning mechanisms for complex robotic manipulation. However, existing approaches share a critical limitation: whether employing explicit linguistic reasoning — which suffers from latency and discretization issues — or utilizing more expressive continuous latent reasoning, they are predominantly confined to static imitation learning, limiting adaptability and generalization. While online reinforcement learning (RL) has been introduced to VLAs to enable trial-and-error exploration, current methods exclusively optimize the vanilla action space, bypassing the underlying physical reasoning process.

LaST-R1 is a unified VLA framework that integrates latent Chain-of-Thought (CoT) reasoning over physical dynamics before action execution, supported by a dedicated RL post-training paradigm:

1. Latent-to-Action Policy Optimization (LAPO): a novel RL algorithm that jointly optimizes the latent reasoning process and action generation. By connecting reasoning with control, it improves representations of physical-world modeling and enhances robustness in interactive environments. 2. Adaptive latent CoT mechanism: allows the policy to dynamically adjust reasoning depth according to environmental complexity.

Results

  • LIBERO benchmark: near-perfect 99.8% average success rate, requiring only a single supervised warm-up stage; convergence speed and performance significantly exceed prior state-of-the-art methods.
  • Real-world deployment: LAPO post-training improves over the initial policy by up to 44% on four complex tasks, including single-arm and dual-arm setups.
  • Links

  • arXiv: https://arxiv.org/abs/2604.28192
*Auto-collected on 2026-05-03*

Tags

#vla#robotics#reinforcement-learning#latent-reasoning#chain-of-thought#manipulation#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619084