[论文] The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents
论文概要
研究领域: ML 作者: Darshan Tank, Baran Nama 发布时间: 2026-07-24 arXiv: 2607.22520
中文摘要
为LLM智能体添加程序性技能通常通过任务成功率的平均提升来评估。然而,这一指标隐藏了一个重要代价:技能也可能使智能体变得更差。本文通过在两个办公自动化基准和三个模型工具链上的近6000次运行中比较有技能和没有技能的智能体来衡量两面性。这使我们能够区分两种结果。回归是指没有技能时能解决但添加技能后失败的任务。残余失败是指有技能和没有技能都失败的任务。本文发现回归足够显著,以至于表现最好的技能之所以优于其他技能,主要是因为回归更少,而非收益更多。本文识别了三种回归原因:(i)技能描述渗透,技能仅因存在于上下文中就改变智能体行为,即使从未被调用;(ii)基础位移,技能规定的过程覆盖了智能体解释输入的方式;(iii)验证位移,过程抑制了智能体本应对其输出执行的检查。分析持续失败揭示了相同的基本模式。现有技能过度强调程序性指导——这一阶段最少导致失败——同时低估了基础和验证,而基础和验证是剩余错误的主要来源。在纠正评估伪影和研究轨迹后,本文发现许多回归和持续失败可以通过更好的基础和验证来恢复。程序性技能应通过将其净效应分解为收益和回归来评估,而非仅靠总体改进。
原文摘要
Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but failed after skills are added. A residual failure is a task that fails both with and without skills. We find that regressions are substantial enough that the best performing skills outperform others primarily by regressing less, not by gaining more. We identify three causes of regression: (i) skill description osmosis, a skill changes an agent's behavior ...
--- *自动采集于 2026-07-28*
#论文 #arXiv #ML #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens