English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Preferences

Forum topic · 小凯 · 2026-05-10

Summary

A 2026 arXiv paper (2605.06656) by Jai Moondra, Ayela Chughtai, Bhargavi Lanka, and Swati Gupta analyzes roughly 89,000 pairwise human comparisons of 52 LLMs across 116 languages in the Arena, showing that best-fitting global Bradley-Terry (BT) rankings are misleading. Nearly two-thirds of decisive votes cancel out, and even the top-50 models by global BT ranking are statistically indistinguishable, with pairwise win probabilities capped at about 0.53. The authors attribute this to strong, structured preference heterogeneity across languages, tasks, and time: grouping votes by language or language family increases agreement sharply and yields ELO score distributions up to two orders of magnitude wider. They introduce the (λ, ν)-portfolio framework—small model sets whose prediction error is at most λ and which cover at least ν of users—formalized as a set-covering variant with guarantees via VC dimension. On Arena data, the algorithm recovers just 5 distinct BT rankings covering over 96% of votes at moderate λ, versus 21% for a global ranking, and a 6-LLM portfolio covers twice as many votes as the global top-6. Portfolios built on the COMPAS dataset with fairness-regularized ensembles can also reveal data blind spots, relevant to policymakers.

Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Preferences

Field: Machine Learning Authors: Jai Moondra, Ayela Chughtai, Bhargavi Lanka, Swati Gupta Published: 2026-05-07 arXiv: 2605.06656

Overview

Ranking LLMs via pairwise human feedback underpins today's leaderboards for open-ended tasks such as creative writing and problem solving. Analyzing ~89,000 comparisons of 52 LLMs across 116 languages in the Arena, the authors show that the best-fitting global Bradley-Terry (BT) ranking is misleading.

Key Findings

  • Nearly 2/3 of decisive votes cancel each other out; even the top-50 models under the global BT ranking are statistically indistinguishable (pairwise win probabilities within the top 50 reach at most 0.53).
  • The failure traces to strong, structured preference heterogeneity across languages, tasks, and time.
  • Language plays a key role: grouping votes by language (and language family) sharply improves agreement, producing ELO score distributions up to two orders of magnitude wider — i.e., highly consistent rankings.
  • What looks like global noise is in fact a mixture of coherent but conflicting subpopulations.
  • The (λ, ν)-Portfolio Framework

    To address this heterogeneity, the paper introduces (λ, ν)-portfolios: small sets of models whose prediction error is at most λ and which "cover" at least a fraction ν of users. The problem is formalized as a variant of set cover, with guarantees derived from the VC dimension of the underlying set system.

    Results

  • On Arena data, the algorithm recovers only 5 distinct BT rankings, covering over 96% of votes at moderate λ, versus 21% for the global ranking.
  • A portfolio of 6 LLMs covers twice as many votes as the global top-6 LLMs.
  • On the COMPAS dataset, portfolios built from ensembles of fairness-regularized classifiers can detect blind spots in the data — of independent interest to policymakers.
--- *Source: arXiv:2605.06656, auto-collected 2026-05-10.*

Tags

#llm#leaderboards#bradley-terry#preference-heterogeneity#evaluation#multilingual#fairness#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619693