论文概要
研究领域: ML
作者: Yinling Zhang, Langchen Liu, Dongbin Xiu, Xueyan Zou, Xu Kuang, Mengdi Wang, Shilong Liu
发布时间: 2026-10-07
arXiv: 2610.10513
中文摘要
语言模型智能体越来越多地被要求开展开放式科学研究,但其结果通常根据已知答案、评分标准或语言模型评审来评分——这些都无法判断一个新的科学模型是否有效。ENSO的AI科学考试(SciExam for ENSO)是一个基准测试,要求智能体从真实观测中构建ENSO(厄尔尼诺-南方涛动,年际气候变率的主要模态)的低阶随机模型。在六小时内,智能体处理观测数据,编写自己的诊断工具(随后被冻结),并仅使用这些诊断作为反馈来构建模型。隐藏评分器随后测试模型是否能重现ENSO的统计特征、恢复未观测变量、预测留出年份,并以相同方式对已发表模型进行评分。在12个智能体系统中,6个产生的模型得分高于已发表模型,主要通过更好的重建和预测实现。较强模型的简化形式各自与ENSO冷暖不对称性两种竞争性解释之一兼容——这是一个任务中从未提及的开放争论。在信息变化条件下对最强系统的受控运行表明,其得分并非来自对过时观测记录的记忆,且接收到的信息塑造了其建模方式。SciExam for ENSO因此能在无已知答案的情况下评估智能体研究,结果表明智能体已能构建有竞争力的模型,其结构涉及科学家仍在争论的问题。
原文摘要
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a pu...
自动采集于 2026-10-09
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。