English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Molecular Déjà Vu: Auditing Verbatim Retrieval of Published Values in Frontier LLMs

Forum topic · 小凯 · 2026-09-09

Summary

A new arXiv paper (2609.05381) audits 22 frontier large language models across 12 molecular property regression benchmarks to determine whether their reported accuracy reflects genuine property prediction or verbatim retrieval of published values from training data. The audit finds that verbatim retrieval is widespread but benchmark-specific: on five datasets, more than 50% of the models show verbatim retrieval, while on the remaining datasets it appears only in isolated cases. Experiments run at two reasoning levels reveal that reasoning changes retrieval behavior — identical experiments on the same molecules with the same prompts are flagged 89% more often at the higher reasoning level than at the lowest one. In the most contaminated cases, the authors test methods to interrupt retrieval and find that the strongest models can still recover combinations of transformed SMILES strings with original labels. Notably, suppressing retrieval makes prediction errors of different models converge in relative terms, suggesting that general predictive ability is not determined solely by the amount of memorized numerical values. The work provides a systematic overview of the extent and depth of verbatim retrieval in molecular regression benchmarks.

Paper Overview

  • Field: Machine Learning
  • Authors: Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
  • Published: 2026-09-04
  • arXiv: 2609.05381
  • Key Points

    LLMs are increasingly evaluated on molecular property benchmarks, but accuracy alone cannot distinguish a model that genuinely predicts a property from one that simply retrieves a published number. This paper systematically audits verbatim retrieval.

    Main Findings

  • Scope of the audit: 22 frontier models were tested on 12 molecular property regression benchmarks for verbatim retrieval of published values.
  • Widespread but benchmark-specific contamination: On five datasets, more than 50% of the LLMs show verbatim retrieval; on the remaining datasets, it appears only in isolated cells.
  • Reasoning changes retrieval: The same experiments, on the same molecules and with the same prompts, are flagged 89% more often at the higher reasoning level than at the lowest one.
  • Retrieval is hard to interrupt: In the most contaminated cases, the authors test ways to interrupt retrieval, and find that the strongest models can in some situations still identify combinations of transformed SMILES strings with original labels.
  • Impact on error structure: Suppressing retrieval makes the prediction errors of different models closer in relative terms, whereas their differing use of verbatim retrieval disperses the errors. This suggests that a model's general predictive ability is not determined solely by how many numerical values it has memorized.

Significance

The work provides the first systematic overview of the extent and depth of verbatim retrieval in LLM molecular regression benchmarks, with implications for how molecular property benchmarks should be interpreted and designed.

Original Abstract (excerpt)

> Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than \(50\%\) of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged \(89\%\) more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cas...

---

*Automatically collected on 2026-09-09.*

Tags

#large-language-models#molecular-property-prediction#data-contamination#benchmarks#verbatim-retrieval#machine-learning#ai-evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634654