Key points
This post is a data-driven comparison of the top 10 AI models globally, based on multi-source benchmark data (BenchLM.ai, LM Market Cap, Artificial Analysis, Google DeepMind), with data current to June 9, 2026.
Scoring methodology
- Provisional Overall score = weighted average of 8 benchmark categories:
- Agentic 22% | Coding 20% | Reasoning 17% | Knowledge 12% | Multimodal 12% | Multilingual 7% | Instruction following 5% | Math 5%
- Anthropic dominates the top 3, with Claude Mythos 5 leading coding, agentic, and reasoning categories.
- Gemini 3.1 Pro is the most complete across widely reported benchmarks, leading scientific reasoning (GPQA Diamond) and HLE.
- Qwen3.7 Max stands out on multilingual tasks and competition math at a mid-range price point ($7.5/M output tokens).
- Grok 4.1 offers the best price-performance ratio with a 2M-token context window at $0.5/M output tokens.
- GPT-5.4 Pro is the most expensive listed model ($180/M output tokens) but tops multimodal understanding and computer control (OSWorld 75%).
Overall ranking (Top 10)
| # | Model | Provider | Score | Highlights | |---|-------|----------|-------|------------| | 1 | Claude Mythos 5 | Anthropic | 99 | #1 coding/agentic/reasoning; SWE-bench Pro 80.3%; 1M+ ctx | | 2 | Claude Fable 5 | Anthropic | 96 | SWE-bench Pro 80.0%; MMMU-Pro 92.7%; 1M+ ctx | | 3 | Claude Opus 4.8 | Anthropic | 94 | SWE-bench Pro 69.2%; HumanEval 95.2%; $25/M out | | 4 | Gemini 3.1 Pro | Google | 92 | GPQA Diamond 94.3%; strongest native multimodal; $12/M out | | 5 | Qwen3.7 Max | Alibaba | 91 | HMMT 97.1% math; #1 multilingual; $7.5/M out | | 6 | GPT-5.4 Pro | OpenAI | 91 | MMMU-Pro 94%; OSWorld 75%; $180/M out | | 7 | GPT-5.5 | OpenAI | 90 | SWE-bench Pro 58.6%; Terminal-Bench 82.7%; $30/M out | | 8 | Gemini 3 Pro Deep Think | Google | 90 | Deep Think reasoning; MATH 89.7%; 2M ctx | | 9 | Grok 4.1 | xAI | 89 | 2M context; best value (8.3/10); $0.5/M out | | 10 | GPT-5.4 | OpenAI | 88 | Balanced (#2 across all categories); OSWorld 75%; $15/M out |
Benchmark table (available scores)
| Model | SWE-bench Pro | HumanEval | MMMU-Pro | GPQA Diamond | HLE | MATH | SimpleBench | |-------|--------------|-----------|----------|--------------|-----|------|-------------| | Claude Mythos 5 | 80.3% | — | 92.7% | — | — | — | — | | Claude Fable 5 | 80.0% | — | 92.7% | — | — | — | — | | Claude Opus 4.8 | 69.2% | 95.2% | — | — | — | — | — | | Gemini 3.1 Pro | 54.2% | 91.6% | 80.5% | 94.3% | 44.4% | 89.7% | 87.4% | | Qwen3.7 Max | 60.6% | — | 79% | 92.3% | 41.4% | 97.1% (HMMT) | — | | GPT-5.4 Pro | — | — | 94% | — | — | — | — | | GPT-5.5 | 58.6% | — | 81.2% | — | — | — | — |
(Scores not listed in the source post are marked with —.)