[论文] In-Context Robot Learning with VLM Agents
研究领域: CV 作者: Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu…
论文概要
研究领域: CV 作者: Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu 发布时间: 2026-09-16 arXiv: 2609.19138
中文摘要
让机器人像人类一样自如地适应陌生环境,仍是具身智能的"登月"目标。有限的演示集合无法覆盖机器人将遇到的每种任务与情境,因此在部署时从上下文学习的能力对泛化至关重要。然而这种上下文学习(ICL)在很大程度上仍超出现有机器人策略的能力。商用视觉语言模型(VLM)广泛的智能体能力(如 GPT-6 Astra)引出一个重要问题:这些模型能否从演示、示例与交互反馈中学习,再将这些信息转化为可执行、可验证的机器人行为——从新的初始状态出发,无需梯度更新或对任务参数的持久更改?我们提出 GPT-Policy——一个面向上下文机器人学习的通用智能体框架,集成三个组件:保留任务相关视觉转换的上下文编译器、提出机器人-工具动作的 VLM,以及验证并执行每个动作且报告结果的约束控制器。我们通过任务成功率与效率指标、跨模型匹配比较以及受控上下文消融,评估其可靠性与局限。真实机器人试验表明:人类视频演示即使在没有机器人动作标签时也能提高任务完成率;在接触敏感任务上,对齐的动作参考进一步带来增益。这些发现将 GPT-Policy 定位为迈向上下文机器人适应的一步,为将 VLM 的通用能力转化为物理行为提供了实证基础,并阐明了可靠部署必须克服的挑战。
原文摘要
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GP...
*自动采集于 2026-09-18*
#论文 #arXiv #CV #小凯