2026 FIFA World Cup, 39 days, 104 matches, 4,494 predictions. Six of the world's strongest LLMs — Claude, GPT, Gemini, Kimi, GLM, and Seed — received a detailed match briefing before every game and filled out a seven-item prediction card: match outcome, scoreline, total goals, group-stage qualification, and more. No answer existed on the internet at the time of prediction.
This was a carefully designed "leak-free" experiment. Not hindsight analysis — the answers were locked into a frozen archive before kickoff.
The result? The six AIs averaged a 63.9% hit rate.
Sounds good? The problem — simply betting the bookmakers' favorite also yields 63.9%.
An Honest Testing Ground
Forecasting benchmarks have an awkward flaw: most events have already happened, and the answers sit somewhere on the internet. Models may not be "predicting" — they may be "reciting." Researchers call this data leakage.
WorldCup Arena sidesteps this trap entirely. Before each match kicked off, the six models each received a briefing — team form, head-to-head history, injuries, tournament context — then filled in a prediction card. The answer did not yet exist when the question was posed. This isn't post-hoc filtering of leaked questions; the design guarantees there is no answer to leak.
Over 39 days, the frozen archive accumulated 4,494 timestamped predictions — genuinely prospective evidence.
Six Models, One Mold
All six models ran their strongest configurations:
- Claude (claude-opus-4-8 thinking, 24k budget)
- GPT (gpt-5.5 reasoning effort high)
- Gemini (gemini-3.1-pro-preview, thinking budget 24,576)
- Kimi (kimi-k2.6 built-in thinking)
- GLM (glm-5.2 thinking)
- Seed (doubao-seed-2-0-pro reasoning effort high)
- One-sided matches (strong vs. weak teams): models predict well — but anyone could
- The most evenly matched games — precisely those with the richest, most detailed briefings — are where models collapse
- Tournament-level predictions (who advances, who wins) were answered reasonably well
All ran at maximum reasoning modes, all with native web search. The strongest lineup money could buy in August 2026.
Yet their behavior was strikingly uniform:
1. They all did the same thing — back the favorite. The six models averaged 63.9% on match outcomes, exactly the same as betting the bookmakers' pick. In other words, six top AIs combined don't beat one odds table.
2. They are too similar to each other. Inter-model agreement far exceeded agreement with the correct answer. The majority vote of the six offered no boost — because they all voted the same way. This isn't an ensemble of six independent judgments; it's closer to six复读机 — six copies of the same machine.
3. They collectively avoid draws and low scores. In football, draws and low-scoring games are the most common upsets. But the models systematically underestimated draw probabilities and total goals. They tend to give a "plausible-looking" scoreline — 2:1, 3:0 — rather than the boring but frequent 1:1.
4. Score predictions cluster on a "prototype outcome." All six models concentrated their scoreline predictions on a handful of combinations, as if stamped from the same mold. Real football score distributions are far more dispersed.
Where Judgment Is Needed Most, They Collapse
The most counterintuitive finding: accuracy is unrelated to how much information a match has, but related to how lopsided it is.
Between the Six Models: Small Gaps, Chaotic Middle
The leaderboard pattern: stable top and bottom, shuffling middle.
Across all 39 days, the strongest and weakest models kept their positions, but the middle rankings kept changing. And throughout, the score gaps between models stayed narrow.
This means: on "predicting real-world events," the current generation of frontier models hasn't truly differentiated. They share the same blind spots, the same preferences, the same failure modes. This isn't a question of "some models are smarter" — it's that "this generation hasn't learned it yet."
What This Tells Us
WorldCup Arena delivers an uncomfortable conclusion: today's strongest LLMs are, at their core, sophisticated "odds复读机" (odds parrots) when it comes to real-event forecasting.
They can read briefings, generate plausible-sounding analysis, and produce self-consistent post-hoc explanations. But when the answer doesn't yet exist — when genuine judgment rather than recitation is required — they merely tie the humblest strategy: backing the favorite.
More alarming is the collective-agreement finding. Six models running independently produce highly similar predictions — meaning ensembling them yields no diversity benefit. If you hoped "multi-model voting" would improve forecast quality, this experiment says: no, because they cast the same vote.
An Honest Exam, an Honest Answer
WorldCup Arena's value isn't "AI fails" — it's that it provides a genuinely honest exam. 4,494 predictions, 39 days, zero leakage — the cleanest evidence to date on LLM forecasting ability.
Its message isn't "models can't predict," but "the current generation hasn't yet surpassed the simplest baseline in real-event forecasting." That's a baseline problem, not a capability-ceiling problem. Will the next generation break through? Unknown. But at least we now have an honest ruler.
The researchers have released all briefings, schedules, official results, and scoring code. For the next big tournament, anyone can measure their own model with the same ruler.
---
Paper: https://arxiv.org/abs/2608.04008
Code & data: The paper states the frozen archive and scoring code will be released.