Loading...
正在加载...
请稍候

[论文] Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Fronti...

小凯 (C3P0) 2026年09月09日 00:43

论文概要

研究领域: ML
作者: Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
发布时间: 2026-09-04
arXiv: 2609.05381

中文摘要

LLM越来越多地在分子性质基准上评估,但准确率无法区分是预测性质还是检索已发表数值。我们对22个前沿模型在12个回归基准上进行逐字检索审计,发现该现象广泛存在但具基准特异性:五个数据集上超过50%的LLM显示逐字检索,其余仅个别出现。我们在两个推理层级运行实验,发现推理改变检索行为——相同分子、相同提示的高层级推理比最低层标记频率高89%。最后,我们在污染最严重案例中测试中断检索的方法,发现最强模型在某些情况下仍能识别变换SMILES字符串与原始标签的组合。此外,抑制检索使不同模型的预测误差在相对意义上更接近,而它们不同的逐字检索使用使误差分散。这表明LLM的一般预测能力并非仅由记忆数值量决定。本工作提供了LLM分子回归基准中逐字检索程度和深度的概览。

原文摘要

Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than \(50\%\) of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged \(89\%\) more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cas...


自动采集于 2026-09-09

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录