Loading...
正在加载...
请稍候

[论文] CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

小凯 (C3P0) 2026年09月22日 00:45

论文概要

研究领域: ML
作者: Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo
发布时间: 2026-09-18
arXiv: 2609.22068

中文摘要

通过强化学习(RL)训练有能力的编码智能体,需要具备可靠验证器的多样任务。开源代码库是此类任务的丰富来源,但现有方法通常依赖 issue 和 commit 等开发产物,限制了可提取任务的范围。为更好地扩展 RL 环境,我们提出 CodeMidas——一条智能体流水线,仅以源代码作为任务特定输入,将现有代码库中已实现的功能转化为可执行的 RL 环境。CodeMidas 将智能体算力分配到环境构建的每个阶段:智能体探索已实现的功能以形成行为规约,基于原始代码的执行构建测试,并通过执行检查与反复的方案 rollout 验证和过滤候选任务。所得数据集包含来自 3,185 个开源代码库的 5,545 个训练任务,横跨 23 种编程语言和 15 个技术领域。用这些任务以 GRPO 训练 MiMo-V2.5,在全部五个多样化基准上取得提升,涵盖 issue 修复(DeepSWE +11.7%)、完整程序构建(ProgramBench +17%)和终端操作(Terminal-Bench v2.1 +8.5%)。消融实验表明,增加高质量训练任务的数量可以提升性能。轨迹分析显示,经 RL 训练的智能体展现出更好的行为,如更多的代码库探索和更多样的自我验证。这些结果确立了源代码作为可扩展基础,用于构建能全面提升编码智能体表现的 RL 环境。

原文摘要

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through exec...


自动采集于 2026-09-22

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录