[论文] TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 40...
论文概要
研究领域: CV 作者: Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang, Di Xu, Yingying Zhu, Dingkang Liang, Xiang Bai, Han Ding 发布时间: 2026-07-29 arXiv: 2607.27205
中文摘要
视觉-语言-动作(VLA)模型通常采用以LLM为中心的 V→L→A 路径:先将视觉观测投影到大语言模型的表征空间,再解码为机器人动作。虽然有效,但每一步策略调用都带来大量计算和内存开销。本文提出 TurboVLA,一种全新的VLA范式,将传统 V→L→A 路径重构为直接的 V+L→A 映射。TurboVLA 不再用大语言模型作为感知与动作之间的中心接口,而是独立编码视觉观测和语言指令,通过轻量级双向视觉-语言交互直接交换信息,并用紧凑的解码器预测连续动作块。这种简洁设计直接从视觉和语言特征构建任务条件表征,显著降低了VLA推理的计算和内存成本。在LIBERO上,TurboVLA仅用0.2B参数、31.2毫秒推理延迟、0.9 GB显存,在消费级RTX 4090上达到97.7%的平均成功率,匹敌甚至超越参数量大得多的VLA策略。
原文摘要
Vision-language-action (VLA) models commonly adopt an LLM-centric V -> L -> A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V -> L -> A pathway as a direct V + L -> A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks wit...
--- *自动采集于 2026-07-31*
#论文 #arXiv #CV #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens