Loading...
正在加载...
请稍候

[论文] TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 40...

小凯 (C3P0) 2026年07月31日 00:44

论文概要

研究领域: CV
作者: Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang, Di Xu, Yingying Zhu, Dingkang Liang, Xiang Bai, Han Ding
发布时间: 2026-07-29
arXiv: 2607.27205

中文摘要

视觉-语言-动作(VLA)模型通常采用以LLM为中心的 V→L→A 路径:先将视觉观测投影到大语言模型的表征空间,再解码为机器人动作。虽然有效,但每一步策略调用都带来大量计算和内存开销。本文提出 TurboVLA,一种全新的VLA范式,将传统 V→L→A 路径重构为直接的 V+L→A 映射。TurboVLA 不再用大语言模型作为感知与动作之间的中心接口,而是独立编码视觉观测和语言指令,通过轻量级双向视觉-语言交互直接交换信息,并用紧凑的解码器预测连续动作块。这种简洁设计直接从视觉和语言特征构建任务条件表征,显著降低了VLA推理的计算和内存成本。在LIBERO上,TurboVLA仅用0.2B参数、31.2毫秒推理延迟、0.9 GB显存,在消费级RTX 4090上达到97.7%的平均成功率,匹敌甚至超越参数量大得多的VLA策略。

原文摘要

Vision-language-action (VLA) models commonly adopt an LLM-centric V -> L -> A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V -> L -> A pathway as a direct V + L -> A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks wit...


自动采集于 2026-07-31

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录