论文概要
研究领域: NLP
作者: Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang
发布时间: 2026-09-17
arXiv: 2609.20804
中文摘要
编码框架(harness)塑造了自主编码智能体将模型能力转化为长周期软件工程性能的方式,然而现有工作通常将框架作为整体系统评估,各组件的有效性尚不明确。为实现组件级比较,我们使用一个轻量级编码框架研究这一问题——其执行循环固定,同时变化三个组件:规划、动作空间和上下文管理。在 SWE-Bench Verified 和 Terminal-Bench 2.1 上评估四个模型,覆盖 176 个匹配设置,涵盖五种上下文管理策略、四种上下文窗口预算以及规划和动作空间的定向消融。我们发现:(1) 随着上下文窗口预算收紧,上下文管理变得越来越有价值,其大部分收益来自防止上下文溢出失败。(2) 将基于规则的省略置于基于 LLM 的摘要之前,在上下文管理策略中提供最强的整体效率;而使省略内容可恢复则增加了模型很少使用的机制,且未带来精度提升。(3) 规划从弱模型的精度脚手架转变为强模型的成本节省器,精度变化不大。(4) 预定义工具改善了 bash 能力较弱模型的性能,而具备 bash 能力的模型可以仅使用 bash 接口有效运行,并实现显著更低的成本,尤其是在以命令行为中心的任务上。轨迹级分析解释了这些效应:上下文管理延长执行轨迹而不实质改变智能体行为,规划改变轨迹的终止位置,动作空间改变代码编写的粒度。这些发现为模型和预算感知的框架设计提供了参考,并提供了一个评估未来框架组件的模块化框架。
原文摘要
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window bud...
自动采集于 2026-09-19
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。