ExecCritic: Learn to Test, Test to Improve for Coding Agents
Fields: cs.AI, cs.CL, cs.SE Authors: Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao arXiv: 2609.09133
Abstract
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence.
ExecCritic combines a test-verify-revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair:
- A Test agent independently generates repository-native tests.
- A fail-closed harness qualifies and freezes the tests.
- A Repair agent revises source code from the tests' execution feedback without changing the tests.
- Holding the base Repair agent fixed, tests from the base Test agent reduce the resolved rate from a no-test baseline of 61.2% to 57.3%.
- Tests from GPT-5.6-sol raise the resolved rate to 65.3%.
- Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%.
- Composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline — without stronger-model or oracle feedback at evaluation time.
Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In *Learn to Test*, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In *Test to Improve*, the Repair agent learns both direct task resolution and feedback-guided revision.
Key Results
On SWE-bench Verified, test quality determines whether feedback helps: