[论文] ExecCritic: Learn to Test, Test to Improve for Coding Agents

论文概要 研究领域: cs.AI, cs.CL, cs.SE 作者: Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao 发布时间: 2…

论文概要

研究领域: cs.AI, cs.CL, cs.SE 作者: Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao 发布时间: 2026-09-08 arXiv: 2609.09133

中文摘要

执行反馈可以引导编码智能体走向正确的仓库修复,但前提是测试能捕捉到问题所要求的行为。智能体生成的测试可能编码不完整或错误的行为目标;当同一条轨迹同时编写补丁和测试时,它们的错误可能一致并产生虚假信心。我们引入ExecCritic,将测试-验证-修订脚手架与针对角色的强化学习配方相结合,用于在其内部训练智能体。该脚手架将测试构建与源代码修复分离:测试智能体独立生成仓库原生测试,一个故障封闭脚手架对测试进行限定并冻结,修复智能体从测试的执行反馈中修订源代码而不改变测试。两个角色都使用Qwen-3.5-35B-A3B作为骨干并分别训练。在'学会测试'中,测试智能体学习生成能区分正确与错误补丁的行为有效测试。在'通过测试改进'中,修复智能体学习直接任务解决和反馈引导的修订。在SWE-bench Verified上,测试质量决定反馈是否有帮助:在固定基础修复智能体的情况下,基础测试智能体的测试将解决率从无测试基线的61.2%降至57.3%,而GPT-5.6-sol的测试将其提升至65.3%。角色特定的后训练将Qwen测试智能体的Base-to-Gold成功率从22.2%提升至62.2%;组合两个后训练的Qwen智能体达到72.6%,比原始无测试基线提升11.4个百分点,且评估时无需更强模型或Oracle反馈。

原文摘要

Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time.


*自动采集于 2026-09-10*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens