English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Preferences

Forum topic · 小凯 · 2026-05-09

Summary

A new arXiv paper (2505.03480) by Jai Moondra, Ayela Chughtai, and Bhargavi Lanka argues that global Bradley-Terry (BT) rankings on LLM leaderboards are misleading. Analyzing ~89K pairwise human comparisons across 116 languages from 52 LLMs on Arena, the authors find that nearly two-thirds of decisive votes cancel out and even the top 50 models are statistically indistinguishable, with pairwise win probabilities capped at 0.53. The failure stems from strong, structured heterogeneity of opinions across language, task, and time. Grouping votes by language massively increases agreement, yielding ELO score distributions two orders of magnitude wider. The paper introduces a (lambda, nu)-portfolio framework: small ensembles of BT rankings achieving at most lambda prediction error while covering at least nu of users. Their algorithm recovers only 5 distinct BT rankings covering over 96% of votes, versus 21% for a single global ranking, and a 6-LLM portfolio covers twice as many votes as the global top 6. Applications to COMPAS-style fair classification are also shown.

Paper Overview

Field: Machine Learning Authors: Jai Moondra, Ayela Chughtai, Bhargavi Lanka Published: 2025-05-09 arXiv: 2505.03480

Abstract (from the paper)

Ranking LLMs via pairwise human feedback underpins current leaderboards for open-ended tasks, such as creative writing and problem-solving. We analyze ~89K comparisons in 116 languages from 52 LLMs from Arena, and show that the best-fit global Bradley-Terry (BT) ranking is misleading. Nearly 2/3 of the decisive votes cancel out, and even the top 50 models according to the global BT ranking are statistically indistinguishable (pairwise win probabilities are at most 0.53 within the top 50 models). We trace this failure to strong, structured heterogeneity of opinions across language, task, and time. Moreover, we find an important characteristic - language plays a key role. Grouping by language (and families) increases the agreement of votes massively, resulting in two orders of magnitude higher ELO score spread - i.e., highly consistent rankings. What appears to be global noise is actually a mixture of coherent but conflicting subpopulations.

To address this heterogeneity in supervised machine learning, we introduce the (λ, ν)-portfolio framework: small ensembles of models that achieve at most λ prediction error while covering at least ν of users. We formalize this as a variant of the set cover problem and provide theoretical guarantees via the VC dimension of the underlying set system. On Arena data, our algorithm recovers just 5 distinct BT rankings that cover over 96% of votes at moderate λ, while a single global ranking covers only 21%. We also provide a portfolio of 6 LLMs that covers twice as many votes as the global top-6 LLMs. Finally, we build portfolios for classification problems using ensembles of fairly regularized classifiers on the COMPAS dataset, showing these portfolios can detect blind spots in data — potentially valuable for policymakers in its own right.

Key Takeaways

  • Global Bradley-Terry leaderboards collapse strong, structured heterogeneity (language, task, time) into a single misleading ranking.
  • Language grouping dramatically increases vote agreement; "global noise" is actually coherent conflicting subpopulations.
  • A portfolio of just 5 BT rankings covers >96% of Arena votes vs. 21% for one global ranking.
  • The portfolio framework generalizes to classification (e.g., fairness-aware COMPAS ensembles) and can reveal data blind spots.
--- *Auto-collected on 2026-05-09*

Tags

#llm#leaderboards#bradley-terry#evaluation#human-feedback#machine-learning#portfolio-methods#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619666