English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration

Forum topic · 小凯 · 2026-05-29

Summary

A paper by Hankyeol Kim and Pilsung Kang (arXiv 2605.27752) examines how evaluation protocol choices affect conclusions in LLM confidence calibration research. Calibration is commonly assessed by comparing two signals: token-probability scores and verbalized confidence, yet these comparisons rest on protocol decisions that are rarely scrutinized. The authors show that calibration conclusions depend heavily on how questions are asked, how answers are elicited, how confidence is scored, and how instances are aggregated. Experiments across eight models and two tasks demonstrate that varying these protocol dimensions changes which signal appears better calibrated and can even reverse the direction of observed gaps—for instance, switching from temperature-scaled to raw token probabilities can flip which model is judged superior. The paper introduces a systematic decomposition of protocol sensitivity and identifies the most consequential dimensions. The findings suggest that current calibration evaluations may be reporting protocol artifacts rather than intrinsic model properties, and that fair comparison requires standardized protocols or sensitivity-aware reporting.

Paper Overview

Field: LLM Authors: Hankyeol Kim, Pilsung Kang arXiv: 2605.27752

Abstract

LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence. These signals are sometimes treated as direct readouts of model uncertainty, but their comparison depends on multiple protocol choices that are rarely examined. The authors show that calibration conclusions are highly sensitive to how questions are asked, how answers are elicited, how confidences are scored, and how instances are aggregated.

Key Findings

  • Across eight models and two tasks, varying protocol dimensions changes which signal appears better calibrated and can even reverse the direction of observed gaps.
  • Example: switching from temperature-scaled to raw token probabilities can flip which model is considered superior.
  • The paper introduces a systematic decomposition of protocol sensitivity and identifies the most consequential dimensions.
  • Current calibration evaluations may be reporting protocol artifacts rather than intrinsic model properties; fair comparison requires standardized protocols or sensitivity-aware reporting.

Discussion

This work is a useful reminder that benchmark conclusions in LLM research often hinge on evaluation setup choices. For practitioners comparing verbalized confidence against token probabilities, the takeaway is to test robustness across protocol variants before drawing conclusions about which model or signal is better calibrated.

--- *Auto-collected on 2026-05-29.*

Tags

#llm#confidence-calibration#token-probability#verbalized-confidence#evaluation-protocols#arxiv#benchmarking

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980500