论文概要
研究领域: NLP
作者: Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
发布时间: 2026-09-21
arXiv: 2609.24972
中文摘要
LLM 智能体的能力很大程度上由其 harness(即围绕冻结骨干模型的提示、控制流、工具、记忆和上下文管理)所放大。最近的方法 increasingly 通过迭代提出和选择智能体 harness 的组件级编辑来自动化这一过程,实际上在智能体系统层面建立了一种递归自我改进(RSI)形式。然而,这种递归演化可能因记忆训练任务而过拟合:在分布内获得较大提升,但在分布外基准上提升缩小甚至消失。我们提出 RRSI(正则化智能体 harness 递归自我改进),将正则化原则引入 harness 自我改进,对演化候选的提出和选择进行约束。提出者采用时间退火预算,限制候选可捆绑的编辑数量,并基于演化历史鼓励探索未走过的路径。选择器配备评论器和剪枝器:评论器筛选特定于基准的提案,剪枝器移除过小、过昂贵或不再有用的更改。这些约束共同偏好可复用的智能体机制而非特定于基准的机制或噪声。在横跨编程、智能体工作空间和工程设计任务的八个基准上,RRSI 在演化针对的分割上提升高达 14.1 分,在五个分布外基准上提升高达 4.7 分,同时产生的 harness 比未正则化演化减少 30% 的策略 token 消耗。代码开源:https://github.com/google-research/rrsi
原文摘要
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selec...
自动采集于 2026-09-23
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。