English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

In-Context Robot Learning with VLM Agents: Introducing GPT-Policy

Forum topic · 小凯 · 2026-09-18

Summary

A forum post introduces the paper 'In-Context Robot Learning with VLM Agents' (arXiv:2609.19138), which proposes GPT-Policy, a general agent framework for in-context robot learning (ICL). The framework addresses the limitation that existing robotic policies cannot learn new tasks at deployment time from context alone. GPT-Policy integrates three components: a context compiler that preserves task-relevant visual transformations, a vision-language model (VLM) that proposes robot-tool actions, and a constrained controller that validates and executes each action while reporting outcomes—enabling adaptation without gradient updates or persistent task-parameter changes. Evaluations use task success rates, efficiency metrics, cross-model comparisons, and controlled context ablations. Real-robot experiments show that human video demonstrations improve task completion even without robot action labels, and aligned action references provide further gains on contact-sensitive tasks. The authors position GPT-Policy as a step toward in-context robot adaptation, providing an empirical basis for translating general VLM capabilities into physical behavior.

Paper Overview

Research area: Computer Vision / Embodied AI Authors: Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu Published: 2026-09-16 arXiv: 2609.19138

Summary

Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies.

The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state—without gradient updates or persistent changes to task-specific parameters?

GPT-Policy Framework

The authors introduce GPT-Policy, a general agent framework for in-context robot learning, integrating three components:

  • Context compiler: preserves task-relevant visual transformations when constructing the model's context.
  • VLM agent: proposes robot-tool actions based on demonstrations, examples, and interaction feedback.
  • Constrained controller: validates and executes each proposed action and reports the resulting outcome.
  • Evaluation and Findings

    Reliability and limitations are assessed through task success rates and efficiency metrics, cross-model matched comparisons, and controlled context ablations. Key findings from real-robot trials:

  • Human video demonstrations improve task completion rates even when no robot action labels are available.
  • On contact-sensitive tasks, aligned action references provide additional gains.
These findings position GPT-Policy as a step toward in-context robot adaptation, offer an empirical basis for translating general VLM capabilities into physical behavior, and clarify the challenges that reliable deployment must overcome.

---

*Auto-collected on 2026-09-18*

Tags

#robotics#vision-language-models#in-context-learning#embodied-ai#gpt-policy#paper#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634943