Key points
- The paradox: Claude Fable 5 scores 80.3% on SWE-Bench Pro, 88.0% on Terminal-Bench 2.1, 85.0% on OSWorld-Verified, and 64.5% on Humanity's Last Exam (with tools) — yet achieves a near-0% full pass rate on Agents' Last Exam (ALE).
- Not just one model: Of 27 top AI configurations tested on ALE's 1,490+ real workflow tasks, 24 had a 0% full pass rate. The paper states: "across mainstream harness and backbone configurations, the average full pass rate is below 1%" (later updated to 2.6%).
- The strongest configuration (Codex + GPT-5.5) scores 82% on Terminal-Bench but under 50% on ALE's easiest tier and under 10% on the hardest.
- Paper: Agents' Last Exam (ALE), arXiv:2606.05405
- Institution: UC Berkeley RDI, with 250+ industry experts
- Scale: 55 subfields, 13 industry clusters, 1,490+ real task instances
- Website: https://agents-last-exam.org/ (leaderboard: https://agents-last-exam.org/leaderboard)
- Scaffold effects: Fable 5's 80.3% on SWE-Bench Pro was achieved with Anthropic's own scaffold. On Scale AI's standardized SEAL leaderboard, GPT-5.4 xHigh scored 59.1% and Claude Opus 4.6 thinking 51.9% — the same models can differ by 20–30 points depending on scaffolding.
- Berkeley's own exploit: An April 2026 Berkeley RDI paper showed single-character modifications could yield 100% scores on some top benchmarks without solving the tasks — evidence many benchmarks measure pattern matching, not capability.
- OSWorld: Considered the hardest to game (~72–75% human baseline, 82% by the Coasty agent), but covers only GUI operations, not cross-tool, multi-file, long-horizon workflows. ALE is effectively a superset of OSWorld + Terminal-Bench plus cross-tool orchestration and end-to-end delivery.
- Long-horizon capability gaps: Existing benchmarks test single-point tasks (fixing a bug, brief desktop operations). ALE tasks require hours-to-weeks of work, multi-software collaboration (CLI + GUI + browser + professional tools), professional judgment, and complex deliverables.
- Cross-tool orchestration: The Generalist Computer-Use Agent must combine visual perception, code execution, tool use, and long-horizon planning. Each step looks fine individually, but chains collapse — the seams between steps are the fatal bottleneck.
- Domain expertise: Civil engineering safety, accounting rules, video color matching — knowledge that general coding ability cannot cover. ALE tasks deliberately embed such judgments.
- For developers: Long-horizon, cross-tool, domain-knowledge integration is the bottleneck; agent architecture (planning, memory, tool use, error recovery) matters more than raw model capability.
- For enterprises: Benchmark scores are reference only — run POCs on real workflows; prioritize long-horizon stability over single-point success rates.
- For researchers: ALE is a living benchmark; 13 of 55 subfields are entirely uncovered by existing benchmarks (Table 1 in the paper). Law, finance, and life sciences may be the next frontiers.
- Sun, Yiyou et al. "Agents' Last Exam." arXiv:2606.05405 (2026). UC Berkeley RDI.
- SWE-Bench Pro: https://www.swe-bench.com/
- OSWorld: https://os-world.github.io/
- Terminal-Bench: https://terminal-bench.com/
- Berkeley RDI "How We Broke Top AI Agent Benchmarks" (April 2026)
- Anthropic Claude Fable 5 Launch (June 9, 2026)
- Scale AI SEAL Leaderboard: https://scale.com/leaderboard
Benchmark details
Why existing benchmarks fall short
The paper argues current benchmarks make unacceptable trade-offs among three dimensions:
1. Realism — Simplified environments or pure Q&A vs. ALE's real VMs, real professional software, and real data files. 2. Coverage — Single domains (code, web, desktop) vs. ALE's 55 subfields across 13 industry clusters, based on the US SOC 2018 occupational classification. 3. Verifiability — Subjective human judgment vs. ALE's structured deliverables with deterministic auto-grading.
Evidence of benchmark inflation
What ALE measures
Tasks must pass three gates:
1. Representativeness: match real professional practice using the actual software experts use (e.g., compositing a running cheetah into a race video — not a single filter operation). 2. Complexity: end-to-end deliverables requiring hours to weeks of expert effort — workflows, not operations. 3. Verifiability: deterministic checking via reference output comparison or explicit rubrics (e.g., replicating an RPG in RPGMaker XP with auto-verifiable map geometry, character stats, and event states).
Tasks are not crowdsourced: experts submit real completed projects, which pass through expert submission → initial screening → engineering implementation → engineer dry-run → expert committee review.
Why agents fail on real workflows
Core insight: benchmark success ≠ GDP impact
The paper's diagnosis: "AI systems have cleared one celebrated benchmark after another... Yet by the metric that ultimately matters, economic output, the broader impact has remained surprisingly muted." ALE is "intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP-relevant impact."
Implications
Limitations
1. No physical-world tasks (robots, physical manufacturing operations). 2. Objective grading remains hard for some domains (creative writing, design aesthetics). 3. Static sandbox — real environments change (software updates, rule changes). 4. No safety-boundary evaluation.
Conclusion
Claude Fable 5's 80% on SWE-Bench vs. near-0% on ALE doesn't mean the model got worse — it means ALE measures completing whole jobs, not fixing single bugs. The next AI milestone isn't +5 points on a benchmark; it's moving full pass rates from 0% to a meaningful number.