[论文] Rules to Tools: Executable Checks for LLM Agents in Scientific Computi...

研究领域: ML 作者: Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke 发布时间: 2026-10-05 arXiv: 2610.00313

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke 发布时间: 2026-10-05 arXiv: 2610.00313

中文摘要

科学编码智能体以书面形式接收方程、边界条件和输出要求,然后须评估其修改的程序。Rules to Tools(R2T)为公开科学要求提供现成可执行检查。匹配的 SciCode 修复组共享书面检查、起始程序、模型和预算;工具组额外收到可调用实现。两个任务 ID 队列中完整修复率:书面 26/30,检查 29/30;三个任务偏向工具,一个偏向书面,十一个打平。八任务队列为 13/16 对 15/16,任务聚类 bootstrap 95% 区间 [-12.5, 43.75] 个百分点。更大共享定义队列两组打平各 13/24。五个开发暴露且带备选起始程序的任务为 3/10 对 7/10。工具组偏向任务 17、77、11;初始检查标记任务 17,任务 77 和 11 报告无违规。任务 37 偏向书面且无初始违规。源码直通 Python 新分支也达 15/16,与专用命令总量一致。匹配 PDE 比较:详细书面 23/24,检查 24/24,检查组模型输出报告值低 31.2%。智能体端节省因队列而异,公开 CPU 使用两队列均升。这些结果测量任务相关修复结果和智能体端成本。

原文摘要

Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference. The larger shared-definition SciCode cohort ties at 13/24 per group. Five development-exposed tasks with alternate s...


*自动采集于 2026-10-05*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens