论文概要
研究领域: ML
作者: Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke
发布时间: 2026-10-05
arXiv: 2610.00313
中文摘要
科学编码智能体以书面形式接收方程、边界条件和输出要求,然后须评估其修改的程序。Rules to Tools(R2T)为公开科学要求提供现成可执行检查。匹配的 SciCode 修复组共享书面检查、起始程序、模型和预算;工具组额外收到可调用实现。两个任务 ID 队列中完整修复率:书面 26/30,检查 29/30;三个任务偏向工具,一个偏向书面,十一个打平。八任务队列为 13/16 对 15/16,任务聚类 bootstrap 95% 区间 [-12.5, 43.75] 个百分点。更大共享定义队列两组打平各 13/24。五个开发暴露且带备选起始程序的任务为 3/10 对 7/10。工具组偏向任务 17、77、11;初始检查标记任务 17,任务 77 和 11 报告无违规。任务 37 偏向书面且无初始违规。源码直通 Python 新分支也达 15/16,与专用命令总量一致。匹配 PDE 比较:详细书面 23/24,检查 24/24,检查组模型输出报告值低 31.2%。智能体端节省因队列而异,公开 CPU 使用两队列均升。这些结果测量任务相关修复结果和智能体端成本。
原文摘要
Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference. The larger shared-definition SciCode cohort ties at 13/24 per group. Five development-exposed tasks with alternate s...
自动采集于 2026-10-05
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。