Overview
Field: LLM Authors: Hankyeol Kim, Pilsung Kang Published: 2026-05-28 arXiv: 2605.27752
Abstract
LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence. These signals are sometimes treated as direct readouts of model uncertainty, but their comparison depends on multiple protocol choices that are rarely examined.
The authors show that calibration conclusions are highly sensitive to:
- How questions are asked
- How answers are elicited
- How confidences are scored
- How instances are aggregated
- A systematic decomposition of protocol sensitivity in calibration evaluation
- Identification of the most consequential protocol dimensions
- Evidence that current calibration evaluations may be reporting protocol artifacts rather than intrinsic model properties
Across eight models and two tasks, varying these protocol dimensions changes which signal appears better calibrated and even reverses the direction of observed gaps. For example, switching from temperature-scaled to raw token probabilities can flip which model is considered superior.
Key Contributions
Implications
The results suggest that fair comparisons of LLM confidence calibration require standardized evaluation protocols, or at minimum, sensitivity-aware reporting that discloses how conclusions vary across protocol choices.
--- *Auto-collected on 2026-05-29*