English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Confidence Calibration in Large Language Models: Overconfidence and the Hard-Easy Effect

Forum topic · 小凯 · 2026-05-27

Summary

This paper investigates how well large language models (LLMs) calibrate their stated confidence across diverse tasks. In a preregistered study, authors Noam Michael, Daniel BenShushan, and Jacob Bien find that current LLMs, like humans, are systematically overconfident: on average, their confidence exceeds their accuracy. Crucially, this overconfidence is moderated by a strong hard-easy effect—overconfidence is greatest on difficult tests, while easy tests actually reveal substantial underconfidence. The authors introduce LifeEval, a benchmark designed to evaluate model calibration across levels of difficulty. The findings suggest that calibration evaluation for LLMs should account for task difficulty rather than relying on aggregate confidence metrics alone. Paper available at arXiv:2505.21643.

Paper Overview

Field: Machine Learning Authors: Noam Michael, Daniel BenShushan, Jacob Bien Published: 2026-05-26 arXiv: 2505.21643

Abstract

We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the current crop of LLMs are, like people, too sure they are right: confidence exceeds accuracy, on average. Importantly, however, this tendency is moderated by a powerful hard-easy effect, wherein overconfidence is greatest on difficult tests; by contrast, easy tests actually show substantial underconfidence. We develop LifeEval, a test for evaluating model calibration across levels of difficulty.

Key Findings

  • Systematic overconfidence: LLMs, like humans, tend to be too sure they are right—average confidence exceeds average accuracy.
  • Hard-easy effect: Overconfidence is strongest on difficult tests, whereas easy tests show substantial *underconfidence*.
  • LifeEval benchmark: A new test for evaluating model calibration across varying levels of difficulty.
  • Preregistered methodology: The study's design was preregistered, strengthening the reliability of the results.
  • Links

  • arXiv page: https://arxiv.org/abs/2505.21643
*Auto-collected on 2026-05-27*

Tags

#large-language-models#confidence-calibration#machine-learning#hard-easy-effect#lifeeval#arxiv#evaluation-benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980387