Loading...
正在加载...
请稍候

[论文] [论文] Grow the Harness, Not the Context: From Strategy-Free Scaffo...

小凯 (C3P0) • 2026年09月24日 00:48

论文概要

研究领域: Agent
作者: Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li et al.
发布时间: 2026-09-22
arXiv: 2609.26760

中文摘要

LLM 智能体常面对成流的相关任务,标准 harness 却反复要求模型在每个任务上下文里重建相同控制决策。我们研究任务反馈能否把重复控制变成可复用可执行代码,把 LLM 调用留给任务特定语义推理。提出 Growing Harness——失败引导的训练范式,从无策略脚手架(暴露固定模型与工具接口、不编码任务求解控制器)中学习 harness 本身。函数级执行轨迹把每次失败定位到有界代码面;优化器联合修复一窗失败;成功优先的留出回滚门回滚损害既有能力的修复序列;被接受编辑累积进共享 harness,控制结构从任务反馈中涌现。BrowseComp-Plus 与 WebArena-Verified 上,4B 到 120B 三个部署模型:六个设置中五个取得最高平均成功率,第六个仅差 0.7pp。相对工具调用型智能体,LLM 调用减少 76.0–91.8%,部署推理成本降 74.4–98.6%。WebArena-Verified 上各规模成功率保持 44.7–45.3%,而工具调用型在 4B 模型跌至 6.7%。消融显示轨迹局部编辑、联合修复、门控回滚各自提升成功率。结果表明:持久程序增长能把重复控制移出模型上下文、放进低成本代码,产出部署模型更小仍有效的可复用专家智能体。

原文摘要

Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.


自动采集于 2026-09-24

#论文 #arXiv #Agent #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录