[论文] Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Fronti...
研究领域: ML 作者: Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler 发布时间: 2026-09-04…
论文概要
研究领域: ML 作者: Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler 发布时间: 2026-09-04 arXiv: 2609.05381
中文摘要
LLM越来越多地在分子性质基准上评估,但准确率无法区分是预测性质还是检索已发表数值。我们对22个前沿模型在12个回归基准上进行逐字检索审计,发现该现象广泛存在但具基准特异性:五个数据集上超过50%的LLM显示逐字检索,其余仅个别出现。我们在两个推理层级运行实验,发现推理改变检索行为——相同分子、相同提示的高层级推理比最低层标记频率高89%。最后,我们在污染最严重案例中测试中断检索的方法,发现最强模型在某些情况下仍能识别变换SMILES字符串与原始标签的组合。此外,抑制检索使不同模型的预测误差在相对意义上更接近,而它们不同的逐字检索使用使误差分散。这表明LLM的一般预测能力并非仅由记忆数值量决定。本工作提供了LLM分子回归基准中逐字检索程度和深度的概览。
原文摘要
Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than \(50\%\) of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged \(89\%\) more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cas...
*自动采集于 2026-09-09*
#论文 #arXiv #ML #小凯