[论文] ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive ...

研究领域: NLP 作者: Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li, Zhaochen Yu, Jinchao Ding, Qiang Gao, Pengyu Zhan, Yuntong Zhang, Tian Cheng, Zhenfei Yin, Yin…

论文概要

研究领域: NLP 作者: Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li, Zhaochen Yu, Jinchao Ding, Qiang Gao, Pengyu Zhan, Yuntong Zhang, Tian Cheng, Zhenfei Yin, Yingcheng Wu, Ling Yang 发布时间: 2026-09-15 arXiv: 2609.17523

中文摘要

我们推出并发布了 ScienceBuddy——一个交互式科学研究工作空间,将持续进化的科学智能体带入研究人员的日常工作流程。ScienceBuddy 在辅助研究人员完成科学任务的同时,将研究人员的请求、反馈和执行证据转化为持续学习的任务与评估标准。其核心是递归之中的递归自我改进范式:将工具链进化和模型强化学习耦合在一起——内层递归在模型固定的情况下改进工具链,外层递归在改进后的工具链下训练模型。工具链进化塑造训练经验,模型学习又为工具链自适应创造新机会。我们展示了研究人员交互、工具链优化和模型学习的案例研究,基准测试涵盖四个科学任务族。通过将 ScienceBuddy 作为研究产品发布,我们使这一范式对科学界可用,并向发现智能体(discovery intelligence)迈出一步:科学AI通过与研究人员的持续协作不断进步,并与它所支持的研究共同演进。网站:http://science-buddy.io

原文摘要

We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks and evaluation rubrics for continual learning. At its core is recursive-in-recursive self-improvement, a paradigm that couples harness evolution with model reinforcement learning: the inner recursion improves the harness with the model fixed, while the outer recursion trains the model under the improved harness. Harness evolution shapes training experience, and model learning creates new opportunities for harness adaptation. We present case studies of researcher interaction, ...


*自动采集于 2026-09-17*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens