Paper Overview
Field: Machine Learning Authors: George Perrett, Javae Elliott, Jennifer Hill, Marc Scott Published: 2026-06-09 arXiv: 2606.11166
Summary
LLMs are increasingly described as performing at the level of human experts on knowledge economy tasks. However, these claims are primarily based on benchmark scores that measure *average* performance across standardized datasets. The paper identifies two key limitations of such benchmarks:
- They often measure performance on content directly included in LLM training data (potential contamination).
- They typically do not assess the reliability of LLM performance or the magnitude of errors — qualities that are critically important in high-stakes contexts.
- Human experts outperform the LLM on average across multiple metrics.
- Human submissions also show less performance variability than the LLM's responses.
Method
The authors design a novel benchmarking task: the model must write computer code to complete a data analysis task. They then compare the performance of a frontier LLM against submissions from human experts, explicitly measuring response variance and error magnitude in addition to average performance.
Findings
Implication
Aggregate benchmark performance can be misleading: even a frontier LLM that matches expert averages on standard benchmarks may be less reliable and produce larger errors on novel, realistic tasks. This challenges narratives that LLMs are ready to fully automate expert-level knowledge work.
---
*Auto-collected from zhichai.net on 2026-06-11.*