English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration (arXiv 2605.27752)

Forum topic · 小凯 · 2026-05-29

Summary

A new arXiv paper (2605.27752) by Hankyeol Kim and Pilsung Kang shows that evaluations of LLM confidence calibration are highly sensitive to protocol choices that are rarely examined. Calibration is typically assessed by comparing token-probability scores against verbalized confidence, often treated as direct readouts of model uncertainty. The authors demonstrate that conclusions depend heavily on how questions are asked, how answers are elicited, how confidence is scored, and how instances are aggregated. Across eight models and two tasks, varying these protocol dimensions changes which signal appears better calibrated and can even reverse the direction of observed gaps; for example, switching from temperature-scaled to raw token probabilities can flip which model is judged superior. The paper introduces a systematic decomposition of protocol sensitivity and identifies the most consequential dimensions. The authors conclude that many current calibration results may reflect protocol artifacts rather than intrinsic model properties, and call for standardized protocols or sensitivity-aware reporting in calibration research.

Overview

Field: LLM Authors: Hankyeol Kim, Pilsung Kang Published: 2026-05-28 arXiv: 2605.27752

Abstract

LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence. These signals are sometimes treated as direct readouts of model uncertainty, but their comparison depends on multiple protocol choices that are rarely examined.

The authors show that calibration conclusions are highly sensitive to:

  • How questions are asked
  • How answers are elicited
  • How confidences are scored
  • How instances are aggregated
  • Across eight models and two tasks, varying these protocol dimensions changes which signal appears better calibrated and even reverses the direction of observed gaps. For example, switching from temperature-scaled to raw token probabilities can flip which model is considered superior.

    Key Contributions

  • A systematic decomposition of protocol sensitivity in calibration evaluation
  • Identification of the most consequential protocol dimensions
  • Evidence that current calibration evaluations may be reporting protocol artifacts rather than intrinsic model properties

Implications

The results suggest that fair comparisons of LLM confidence calibration require standardized evaluation protocols, or at minimum, sensitivity-aware reporting that discloses how conclusions vary across protocol choices.

--- *Auto-collected on 2026-05-29*

Tags

#llm#calibration#confidence-estimation#arxiv#evaluation-methodology#research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980522