[论文] Skill-Space Shooting for Autonomous Robot Policy Improvement
研究领域: ML 作者: Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao 发布时间: 2026-09-29 arXiv: 2609.38178
论文概要
研究领域: ML 作者: Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao 发布时间: 2026-09-29 arXiv: 2609.38178
中文摘要
部署在物理世界中的机器人必须能够在遇到新情况和故障时超越初始训练水平持续提升。为了使这种改进能够跨任务扩展,它必须有效利用经验,而不需要人类对每个修正进行演示。近期基于智能体(agentic)的系统通过基础模型自主组合已学习的行为来完成任务,从而减少对人类努力的依赖。然而,以这种方式完成任务本身并不能教会任务策略克服自身故障——这需要将这些行为转化为可学习的策略修正。我们的洞见是:许多此类修正是常见的短行为或技能,它们跨任务反复出现,描述的是基础模型可以从场景中推理出的动作。我们提出技能空间射击(skill-space shooting)方法,利用基础模型引导通过这些可复用技能探索修正方案,并将成功试验转化为策略改进。真实世界实验表明,策略在自主运行中实现了反复改进,同时技能还可以共享,从而减少在新任务上改进所需的教学量。通过将可复用技能作为纠正性监督的来源,技能空间射击实现了任务内和跨任务的可扩展、可泛化策略改进。更多结果和视频见项目页面。
原文摘要
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a sce...
*自动采集于 2026-10-01*
#论文 #arXiv #ML #小凯