论文概要
研究领域: ML
作者: Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi
发布时间: 2026-10-01
arXiv: 2610.02204
中文摘要
在多样化任务上构建可靠的机器人能力需要大量人力来开发和维护技能、设计奖励、整合感知与控制。我们提出"重建、练习、实战"(RPG)框架,无需更新模型权重即可实现机器人执行系统的自主改进。RPG 从离线数据集中识别操作能力,在仿真中构建相关的练习任务。练习过程中,RPG 利用执行反馈、特权模拟器状态和可用数据集视频诊断失败原因,开发新的可复用符号技能、精化现有技能并修订系统提示。跨任务评估在保留合并修订前逐一测试候选变更。测试时,多模态 LLM 使用生成的系统提示和技能库协调感知与机器人控制。在 22 个操作任务的留出初始化上,RPG 将任务成功率从首轮练习后的 28.6% 提升至 15 轮后的 95.0%,超过所有评估的基线,包括 ASPIRE(75.5%)和由 GPT-6 Astra Pro 驱动的 CaP-Agent0(60.0%)。经过通用校准和硬件适配后,冻结系统在所有 30 次物理实验中全部成功。
原文摘要
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reu...
自动采集于 2026-10-03
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。