English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The War on a Hair's Width: A One-Centimeter Ruler Claiming to Measure 0.1 Millimeters

Forum topic · 小凯 · 2026-09-24

Summary

A Chinese tech forum essay examines a paper by Atul Anand, 'Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM' (arXiv:2609.28082), which performs a metrological audit of LLM-as-a-judge evaluation. Analyzing 373,019 judge verdicts through Generalizability Theory, the paper argues that single-judge, point-based scoring has a structural reliability ceiling that more test items cannot fix, because system-by-judge interaction variance is never averaged away. The minimum reliably detectable score difference on 0–5 scales ranges from 0.41 to 1.24 points—while the median claimed improvement in the literature is only 0.28 points. Of 17 published MT-Bench improvement claims, all fell below this noise floor. Switching to pairwise preference comparison with swapped presentation orders drops judge-interaction variance by two orders of magnitude, reaching a generalizability coefficient of 0.986—though position bias (+8.6 percentage points for answers shown first) must be controlled. An audit of 628 arXiv papers found fewer than a quarter disclose how many evaluation runs they performed. The essay concludes that most leaderboard gaps cited in the LLM era may be measurement noise rather than real model differences.

The Problem

LLM leaderboards routinely declare victories by margins of 0.1–0.2 points. But can the measuring instrument itself—LLM-as-a-judge—actually resolve differences that small? On September 24, 2026, Atul Anand published *Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM* (arXiv:2609.28082, 18 pages), a metrological calibration of LLM-based evaluation. The answer: most frequently cited score gaps are smaller than the instrument's own noise.

Since 2023, when MT-Bench showed GPT-4's preferences agreed with human judgments over 80% of the time, LLM-as-a-judge has replaced expensive human evaluation. Prior critiques focused on *bias*—verbosity bias, position bias, self-preference—as if the watch were merely running a few minutes fast. Anand asks a different question: what is the *minimum graduation* on this ruler?

Key Points

Generalizability Theory Applied to LLM Judges

  • 373,019 judge verdicts were decomposed using Generalizability Theory (Cronbach et al., 1972): system variance, item variance, judge variance, and their interactions.
  • The fatal term is the system × judge interaction (σ²sj): a judge who favors verbose answers systematically inflates a verbose model across every item. Adding more items cannot average this away.
  • Asymptotic result: with a single judge, the generalizability coefficient converges to σ²s/(σ²s + σ²sj) as items increase—items saturate; judges do not.
  • This ceiling is a flaw of *point-based rubric scoring* as a measurement format, not of LLM judges specifically; human judges using 1–5 scales hit the same wall.
  • The Noise Floor vs. Claimed Improvements

  • The minimum reliably detectable difference on 0–5 scales is 0.41 to 1.24 points (at native benchmark sizes).
  • The median claimed improvement in the literature: 0.28 points—below even the lowest floor.
  • All 17 published MT-Bench improvement claims fell below the floor. 70% of win-rate claims fall below the noise floor of pairwise comparison itself.
  • Near the ceiling, reliability gains require exponentially more items—explaining why such calibration was never done before.
  • The Fix: Ask Which, Not How Good

  • Replace point scores with pairwise preference (A vs. B, both presentation orders).
  • Result: judge-interaction variance drops below 1% of system variance; a single judge achieves a generalizability coefficient of 0.986 (bootstrap CI [0.934, 1.000], 11 systems). This is a retroactive vindication of AlpacaEval's protocol—provided presentation order is swapped.
  • Cost: position bias. The answer shown first wins by +8.6 percentage points—larger (1.23×) than the median published win-rate claim (~7 pp) across 53 surveyed papers.

Audit of the Field

Anand audited 628 arXiv papers (dual model coding, kappa = 0.73 vs. blind human coding):

1. Fewer than 25% disclose how many times their evaluation was run. 2. Only 46–67% report any uncertainty (error bars, CIs, variance). 3. Most conclusions are therefore unverifiable and incomparable as reported.

Conclusion

The paper's three compressed findings: (1) single-judge point scoring has a structural ceiling no number of items can break; (2) asking *which is better* makes one judge sufficient—if you swap presentation order; (3) most claimed improvements in the current literature are smaller than the instrument's own floor. Before citing a '0.21-point lead,' ask: how many items? How many judges? How many runs?

Reference: Anand, A. (2026). *Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM*. arXiv:2609.28082. https://arxiv.org/abs/2609.28082

Tags

#llm-evaluation#llm-as-a-judge#benchmarking#metascience#measurement-theory#arxiv#mt-bench#statistical-reliability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635167