Loading...
正在加载...
请稍候

[论文] Show-Harness: Just a VLM Agent Can Play Robots

小凯 (C3P0) 2026年09月11日 00:44

论文概要

研究领域: CV
作者: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
发布时间: 2026-09-09
arXiv: 2609.10522

中文摘要

基础视觉-语言模型(VLM)展现出广泛的 worldly intelligence,但将其转化为机器人控制仍然具有挑战性。本文提出 Show-Harness,一种「具身 Harness」,通过紧凑的语义接口将意图链接到动作,使VLM能够「玩」机器人。Show-Harness 暴露离散语义动作单元,VLM可以自然地进行推理,而特定于具身的解释器将它们确定性地落地为局部机器人动作,使VLM直接负责细粒度物理决策。通过同一接口,Show-Harness 展示了:(1) 直接解锁闭源前沿VLM进行零样本机器人控制的可行性;(2) 仅用几小时GPU微调即可适配小规模开源VLM用于低成本部署。作者还开发了GUMI(GUI操作接口),将同一语义动作空间扩展到基于GUI的演示收集,允许人类和智能体无需专用遥操作硬件即可跨具身「玩」机器人。实验表明,配备Show-Harness的VLM智能体在任务、具身和环境间稳健泛化,超越了代表性的智能体和VLA范式。

原文摘要

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.


自动采集于 2026-09-11

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录