The Problem
LLM leaderboards routinely declare victories by margins of 0.1–0.2 points. But can the measuring instrument itself—LLM-as-a-judge—actually resolve differences that small? On September 24, 2026, Atul Anand published *Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM* (arXiv:2609.28082, 18 pages), a metrological calibration of LLM-based evaluation. The answer: most frequently cited score gaps are smaller than the instrument's own noise.
Since 2023, when MT-Bench showed GPT-4's preferences agreed with human judgments over 80% of the time, LLM-as-a-judge has replaced expensive human evaluation. Prior critiques focused on *bias*—verbosity bias, position bias, self-preference—as if the watch were merely running a few minutes fast. Anand asks a different question: what is the *minimum graduation* on this ruler?
Key Points
Generalizability Theory Applied to LLM Judges
- 373,019 judge verdicts were decomposed using Generalizability Theory (Cronbach et al., 1972): system variance, item variance, judge variance, and their interactions.
- The fatal term is the system × judge interaction (σ²sj): a judge who favors verbose answers systematically inflates a verbose model across every item. Adding more items cannot average this away.
- Asymptotic result: with a single judge, the generalizability coefficient converges to σ²s/(σ²s + σ²sj) as items increase—items saturate; judges do not.
- This ceiling is a flaw of *point-based rubric scoring* as a measurement format, not of LLM judges specifically; human judges using 1–5 scales hit the same wall.
- The minimum reliably detectable difference on 0–5 scales is 0.41 to 1.24 points (at native benchmark sizes).
- The median claimed improvement in the literature: 0.28 points—below even the lowest floor.
- All 17 published MT-Bench improvement claims fell below the floor. 70% of win-rate claims fall below the noise floor of pairwise comparison itself.
- Near the ceiling, reliability gains require exponentially more items—explaining why such calibration was never done before.
- Replace point scores with pairwise preference (A vs. B, both presentation orders).
- Result: judge-interaction variance drops below 1% of system variance; a single judge achieves a generalizability coefficient of 0.986 (bootstrap CI [0.934, 1.000], 11 systems). This is a retroactive vindication of AlpacaEval's protocol—provided presentation order is swapped.
- Cost: position bias. The answer shown first wins by +8.6 percentage points—larger (1.23×) than the median published win-rate claim (~7 pp) across 53 surveyed papers.
The Noise Floor vs. Claimed Improvements
The Fix: Ask Which, Not How Good
Audit of the Field
Anand audited 628 arXiv papers (dual model coding, kappa = 0.73 vs. blind human coding):
1. Fewer than 25% disclose how many times their evaluation was run. 2. Only 46–67% report any uncertainty (error bars, CIs, variance). 3. Most conclusions are therefore unverifiable and incomparable as reported.
Conclusion
The paper's three compressed findings: (1) single-judge point scoring has a structural ceiling no number of items can break; (2) asking *which is better* makes one judge sufficient—if you swap presentation order; (3) most claimed improvements in the current literature are smaller than the instrument's own floor. Before citing a '0.21-point lead,' ask: how many items? How many judges? How many runs?
Reference: Anand, A. (2026). *Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM*. arXiv:2609.28082. https://arxiv.org/abs/2609.28082