English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Why 67-Model Ensembles Can't Beat the Single Best LLM: The Overlooked Co-Failure Ceiling

Forum topic · ✨步子哥 · 2026-06-26

Summary

A paper by Josef Chen (KAIKAKU), "When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models," argues that the AI orchestration field has been optimizing the wrong metric. Instead of pairwise error correlation ρ, the decisive quantity is β — the probability that all models in a pool fail simultaneously on the same query. No selection strategy (routing, voting, cascades, Mixture-of-Agents) can exceed an accuracy ceiling of 1 − β, and ρ is mathematically blind to β, mirroring the Gaussian-copula mispricing that underestimated joint defaults in 2008 CDOs (ρ underestimates β by roughly 2.5× at 67-model scale). Across 67 frontier models from 21 vendors, four trained routers captured essentially zero oracle gain, with an LLM-as-router picking the single strongest model 100% of the time. The paper offers a $0 pre-deployment "certificate" (beta_certificate.py): count all-wrong queries, compute a Clopper-Pearson lower bound on β, and derive the maximum possible gain over the best single model. A controlled experiment on GPQA-Diamond shows β flips from ~0 (multiple choice) to 0.127 (open-ended) on the same questions, meaning format, not subject matter, drives co-failure — and benchmark gains may not transfer to open-ended production settings.

Why 67-Model Ensembles Can't Beat the Single Best LLM: The Overlooked Co-Failure Ceiling

The scenario: an AI architect's dilemma in 2026

Imagine you're an AI architect at a company. The 2026 model market looks like a stock exchange: GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro, Grok-4.3, DeepSeek V4, Qwen3.7-Max, Kimi K2.7... 67 frontier models from 21 vendors, priced from $30/M tokens down to $0.1/M, with each new generation turning the last into a bargain.

Your boss asks: "Can we combine these models to be more accurate than any single one?"

You open the literature and find everyone making decisions based on one metric: pairwise error correlation ρ — how correlated models' errors are. Low ρ means models fail on different queries, so combining them should help, just like diversification in portfolio theory.

Sounds reasonable. So you train a router, try majority voting, cascading, Mixture-of-Agents... and discover: the router captures almost no gain, and LLM-as-router (GPT-5-mini) picks the single strongest model on 100% of queries. You suspect the router is too weak, so you try gradient boosting, multiclass predictors, even letting an LLM read all models' strengths before choosing — four routers, none beats the single best model.

What's going wrong?

The core finding: ρ is the wrong metric, β is the right one

Josef Chen (KAIKAKU), in the paper *When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models*, gives the answer: the field has been looking at the wrong number.

The core insight fits in one inequality:

> The accuracy ceiling of any selection strategy (routing, voting, cascades) = 1 − β

where β = the probability that all models fail simultaneously on the same query.

The inequality is trivial: if every model is wrong, any strategy that selects one of their answers must also be wrong. Yet this trivial fact has been overlooked — everyone optimizes ρ while nobody measures β.

The paper proves something sharper: ρ mathematically cannot identify β. There exist error distributions with identical marginal error rates and identical pairwise correlations ρ but completely different β. In other words, ρ is "blind" to β.

The $0 certificate: know whether combination helps before deploying

The most practical contribution is a "free certificate." Before training any router, before paying for labeled data, before any engineering investment — you only need one thing:

Count how many of n queries all models get wrong; call it K.

Then compute a Clopper-Pearson confidence lower bound on β. Plug into 1 − β_lo − a_sb (where a_sb is the single best model's accuracy), and you get a confidence upper bound on the maximum gain any selection strategy can achieve over the single best model.

If that number is below your orchestration overhead (router training cost, extra latency, maintenance), stop — no strategy can pay off.

It's a $0 test: no router training, no pairwise ρ, no complex analysis. Just a labeled test set and the models' answers. The paper open-sources this as beta_certificate.py.

Why ρ "underestimates" β — the ghost of 2008 CDOs

Could you infer β from ρ, say via a single-factor Gaussian copula? The paper says: no — and it systematically underestimates.

At 67-model market scale, a correctly calibrated tetrachoric single-factor model still predicts β about 2.5× lower than observed (on open math, 90% CI 1.7–3.4, k=17). The model says β=0.023; the actual β=0.052. And the underestimation worsens as the pool grows.

If this story feels familiar, it's because it is the mathematical structure of the 2008 subprime crisis. Back then, Gaussian copula models priced CDOs using pairwise default correlations to infer the probability of simultaneous defaults. In normal times the estimates looked fine; in extreme conditions, assets crashed together — tail correlations far exceeded pairwise correlations. The models systematically underestimated tail risk.

The paper explicitly cites this analogy: "body-vs-tail base-correlation smile of Gaussian-copula portfolio-credit (CDO) models." Same math, same trap — just moved from subprime mortgages to language models.

Format decides everything: the same questions, β from 0 to 0.127

The paper's most elegant experiment is a content-controlled one. The authors took the same GPQA-Diamond questions and tested twice:

  • Multiple-choice version: β ≈ 0 (almost no queries where all models fail together)
  • Open-ended version (options removed): β = 0.127 (about 8% of questions missed by all models)
  • Same questions, same models — only the answer format changed. Average accuracy fell from 0.66 to 0.51; the best model fell from 0.91 to 0.77.

    Co-failure is determined not by subject matter but by the "openness" of the answer format. Multiple choice gives models a strong prior — even guessing yields 1/4 odds — crushing β to near zero. Remove the options, and the questions models truly can't solve are exposed — and exposed collectively.

    Practical implication: an orchestration strategy validated on multiple-choice benchmarks may fail in real open-ended settings. Many papers report routing gains on MMLU and ARC, but production looks more like open-ended generation — where β is much higher and ceilings much lower.

    Two kinds of ceilings that ρ cannot distinguish

    Ceiling-bound

  • Where: open math, code generation, open-ended science
  • Signature: β > 0; all models fail together on some queries
  • Consequence: the 1 − β ceiling is real and unbeatable
  • Data: MATH-500 β=0.052; code competitions β=0.079; open GPQA β=0.127
  • Realizability-bound

  • Where: multiple-choice science (GPQA-Diamond MCQ)
  • Signature: β ≈ 0, ceiling high, oracle gain G=0.154 is large
  • Consequence: the ceiling isn't the problem — but deployable routers can't capture the gain
  • Data: oracle reaches 1.000 on GPQA MCQ, but learned routers achieve ~0 gain
  • Both dilemmas imply the same practical outcome: without strong query-level routing signals, model combinations rarely beat the single best model. But the causes differ — one mathematical, one engineering — and ρ can't tell you which one you're in.

    The 67-model experiments: four routers, total failure

    On a 15-model pool, the paper tested four routers:

    | Router | Fraction of G captured | |--------|------------------------| | TF-IDF + logistic regression | 9% (CI [-0.67, 0.50], not significant) | | Gradient boosting (word+char features) | -9% (negative) | | Multiclass best-model predictor | -127% (worse than routing) | | LLM-as-router (GPT-5-mini) | 0% (picks single best 100%) |

    None significantly beats the single best model. Why? The paper's explanation is incisive: "prompt carries little signal about which model will be the one that is right when the frontier disagrees."

    This directly contradicts optimistic results from papers like "More Agents Is All You Need." With quality mismatch, majority-voting gains are negative (-0.10 on hard sets, -0.02 on saturated ones). Mixing unequal models lets weak models' numbers vote down the strong ones.

    But with quality matching, diversity does help

    The paper doesn't dismiss diversity entirely. Under quality matching (all models at similar accuracy), low ρ does yield gains:

  • Self-MoA (repeated sampling from one best model, ρ_intra=0.80) vs. heterogeneous ensembling (6 accuracy-matched models, ρ_inter=0.42)
  • At k=3 independent samples per side, heterogeneous ensembling beats Self-MoA by +0.027
  • Positive in 100% of 60 resamples (direction robust, magnitude partition-dependent)
  • This confirms the diversification theorem: with matched quality, lower pairwise correlation buys larger diversifiable gains. But note the premise — in real deployments, model accuracies are rarely matched. Hence the practical bottom line: naive diversity is a liability.

    Three practical engineering takeaways

    1. Run the $0 certificate before deploying

    Before any router training, count the all-wrong rate K/n and run beta_certificate.py. If the certified maximum gain is below orchestration overhead, walk away.

    2. Monitor β, not ρ

    ρ is misleading. Two model pools can share the same ρ but have completely different β. Measure β directly — it needs only a labeled test set, no pairwise computation.

    3. Watch the format, not the subject

    Is your benchmark multiple-choice or open-ended? That matters more than "math vs. science vs. code." Multiple-choice β≈0 can make your orchestration look effective, but it will fail when transferred to open-ended settings.

    A deeper reflection: AI's "shared blind spots"

    β measures something profound: the blind spot shared by all models. On open math, 5.2% of questions defeat all 67 models from 21 vendors. These aren't "hard" questions — they're products of every model's shared training data, architectures, and alignment.

    The 2008 isomorphism isn't only mathematical. The CDO crisis happened because all rating agencies, banks, and investors shared one assumption (national home prices never fall), so they all underestimated tail risk. Models' shared blind spots are similarly products of "shared assumptions" — possibly common training-data contamination, common alignment preferences, common refusal patterns.

    The paper doesn't develop this analogy, but it implies a warning: as training becomes more homogeneous (everyone uses RLHF, similar data recipes), β may rise. More models doesn't mean more diversity — if failure modes converge, adding models just adds homogeneous redundancy.

    The paper's final line is worth remembering:

    > "On open-ended tasks the best models increasingly fail alike, so the lever is failure-mode dispersion and market churn, not peak capability or model count."

    The real lever isn't stronger models or more models — it's finding models with different failure modes, and waiting for market churn to bring new heterogeneity. This echoes portfolio theory: diversification gains come not from asset count but from low correlation. And when all assets crash together, diversification fails — because β, not ρ, is what decides.

    ---

    Paper: When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

    Author: Josef Chen (KAIKAKU)

    Core tool: beta_certificate.py — a $0 pre-deployment certificate: input the all-wrong count, output an upper bound on any combination strategy's gain

    Experiment scale: 67 frontier models, 21 vendors, ≈$270 reported experiment cost

    Key numbers:

  • MATH-500: β=0.052, ρ underestimates β 2.5×
  • Code competitions: β=0.079, ρ underestimates β 3.1×
  • GPQA MCQ: β≈0 (realizability-bound)
  • GPQA open-ended: β=0.127 (format-flip regime)
  • Learned routers' capture of G: ≈0%
  • LLM-as-router: picks single best model 100% of the time

Tags

#llm-orchestration#model-ensembling#routing#co-failure-ceiling#majority-voting#mixture-of-agents#portfolio-theory#ai-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208157