[论文] EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided...

研究领域: ML 作者: Python Song, Zhixuan Liang, Kelsey Fu, Mengdi Wang, Junfeng Yang, Shilong Liu 发布时间: 2026-10-07 arXiv: 2610.10498

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Python Song, Zhixuan Liang, Kelsey Fu, Mengdi Wang, Junfeng Yang, Shilong Liu 发布时间: 2026-10-07 arXiv: 2610.10498

中文摘要

机器人基础模型提供强大的视觉运动控制,但当物体位置或任务指令变化时性能可能下降。进一步提升通常需要在大量机器人数据上进行后训练,通过遥操作等方法收集数据的成本很高。智能体外壳可以在模型周围进行适配,但当前的自进化外壳在决定 pursue 哪些代码和技能变更时,对机器人试验的使用效率低下。我们引入EmbodiedRSI——一个自进化的智能体外壳,自主决定下一步在哪里探索,并将由此产生的物理交互转化为改进的代码和技能。EmbodiedRSI通过快慢双系统架构实现这一目标——竞争性的代码和技能假设被维护在假设图中。信息价值实验选择能够区分这些假设的物理实验。实验结果指导代码-技能协同进化。慢系统构建层次化记忆,奖励锚定记忆学习根据对快系统后续改进的价值选择有效记忆。在RoboCasa365上,EmbodiedRSI达到77.0%的总体成功率和71.3%的复合未见任务成功率,而最佳基线为40.1%。EmbodiedRSI在LIBERO-Pro上也达到86.8%的总体成功率。除基准性能外,EmbodiedRSI还能零样本迁移到真实机器人,在多个挑战性任务中实现71.3%的总体成功率。

原文摘要

Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-...


*自动采集于 2026-10-09*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens