[论文] CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
论文概要
研究领域: ML 作者: Qijia He, Jiayi Cheng, Chenqian Le 发布时间: 2026-07-22 arXiv: 2507.17084
中文摘要
编程智能体越来越多地在可执行环境中运行,失败的尝试产生可操作的反馈,而非仅仅是错误答案。现有成本感知系统通常将此类失败视为级联决策:先尝试廉价模型,再将困难案例升级到更强更贵的模型。然而,在编程中,执行反馈也使进一步的廉价模型恢复变得有价值,这引出一个预算部署问题:智能体何时应花费更多廉价计算,何时应升级?我们将这种失败后的决策形式化为异构动作上的恢复路由,并从执行轨迹中训练监督路由器。为使同一路由器在变化的预算下可用,我们添加\textbf{保形风险控}(CRC)层,无需重新训练即可选择部署时成本惩罚,并在可交换性下提供边际预期成本控。在五个编程基准的保留失败案例上,廉价恢复和升级展现出互补的成功模式。校准前沿优于固定动作、仅提示路由器和二元级联基线;在主要GPT-5.4-nano/GPT-5.4设置中,一个CRC校准前沿点超过始终升级的解决率,同时使用其平均恢复成本的35%。代码已开源。
原文摘要
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment...
--- *自动采集于 2026-07-23*
#论文 #arXiv #ML #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens