English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agents' Last Exam Deep Dive: Why Claude Fable 5 Scores 80% on SWE-Bench But Near-Zero on Real Workflows

Forum topic · 小凯 · 2026-06-18

Summary

Agents' Last Exam (ALE), a benchmark developed by UC Berkeley RDI with 250+ industry experts (arXiv:2606.05405), tests AI agents on 1,490+ real professional workflow tasks spanning 55 subfields and 13 industry clusters. The headline finding: Claude Fable 5, which scores 80.3% on SWE-Bench Pro and 88% on Terminal-Bench 2.1, achieves a near-0% full pass rate on ALE. In fact, 24 of 27 top AI configurations scored 0% on full pass rates, with the average full pass rate below 1% (later updated to 2.6%). The paper argues existing benchmarks make unacceptable trade-offs among realism, coverage, and verifiability, and that scaffold choice alone can swing SWE-Bench scores by 20-30 points. ALE uses real VMs, professional software, deterministic auto-grading, and end-to-end deliverables. The real bottlenecks for agents are long-horizon planning, cross-tool orchestration, and domain expertise integration rather than single-step skills. Benchmark success, the authors contend, does not equal GDP-relevant economic impact.

Key points

  • The paradox: Claude Fable 5 scores 80.3% on SWE-Bench Pro, 88.0% on Terminal-Bench 2.1, 85.0% on OSWorld-Verified, and 64.5% on Humanity's Last Exam (with tools) — yet achieves a near-0% full pass rate on Agents' Last Exam (ALE).
  • Not just one model: Of 27 top AI configurations tested on ALE's 1,490+ real workflow tasks, 24 had a 0% full pass rate. The paper states: "across mainstream harness and backbone configurations, the average full pass rate is below 1%" (later updated to 2.6%).
  • The strongest configuration (Codex + GPT-5.5) scores 82% on Terminal-Bench but under 50% on ALE's easiest tier and under 10% on the hardest.
  • Benchmark details

  • Paper: Agents' Last Exam (ALE), arXiv:2606.05405
  • Institution: UC Berkeley RDI, with 250+ industry experts
  • Scale: 55 subfields, 13 industry clusters, 1,490+ real task instances
  • Website: https://agents-last-exam.org/ (leaderboard: https://agents-last-exam.org/leaderboard)
  • Why existing benchmarks fall short

    The paper argues current benchmarks make unacceptable trade-offs among three dimensions:

    1. Realism — Simplified environments or pure Q&A vs. ALE's real VMs, real professional software, and real data files. 2. Coverage — Single domains (code, web, desktop) vs. ALE's 55 subfields across 13 industry clusters, based on the US SOC 2018 occupational classification. 3. Verifiability — Subjective human judgment vs. ALE's structured deliverables with deterministic auto-grading.

    Evidence of benchmark inflation

  • Scaffold effects: Fable 5's 80.3% on SWE-Bench Pro was achieved with Anthropic's own scaffold. On Scale AI's standardized SEAL leaderboard, GPT-5.4 xHigh scored 59.1% and Claude Opus 4.6 thinking 51.9% — the same models can differ by 20–30 points depending on scaffolding.
  • Berkeley's own exploit: An April 2026 Berkeley RDI paper showed single-character modifications could yield 100% scores on some top benchmarks without solving the tasks — evidence many benchmarks measure pattern matching, not capability.
  • OSWorld: Considered the hardest to game (~72–75% human baseline, 82% by the Coasty agent), but covers only GUI operations, not cross-tool, multi-file, long-horizon workflows. ALE is effectively a superset of OSWorld + Terminal-Bench plus cross-tool orchestration and end-to-end delivery.
  • What ALE measures

    Tasks must pass three gates:

    1. Representativeness: match real professional practice using the actual software experts use (e.g., compositing a running cheetah into a race video — not a single filter operation). 2. Complexity: end-to-end deliverables requiring hours to weeks of expert effort — workflows, not operations. 3. Verifiability: deterministic checking via reference output comparison or explicit rubrics (e.g., replicating an RPG in RPGMaker XP with auto-verifiable map geometry, character stats, and event states).

    Tasks are not crowdsourced: experts submit real completed projects, which pass through expert submission → initial screening → engineering implementation → engineer dry-run → expert committee review.

    Why agents fail on real workflows

  • Long-horizon capability gaps: Existing benchmarks test single-point tasks (fixing a bug, brief desktop operations). ALE tasks require hours-to-weeks of work, multi-software collaboration (CLI + GUI + browser + professional tools), professional judgment, and complex deliverables.
  • Cross-tool orchestration: The Generalist Computer-Use Agent must combine visual perception, code execution, tool use, and long-horizon planning. Each step looks fine individually, but chains collapse — the seams between steps are the fatal bottleneck.
  • Domain expertise: Civil engineering safety, accounting rules, video color matching — knowledge that general coding ability cannot cover. ALE tasks deliberately embed such judgments.
  • Core insight: benchmark success ≠ GDP impact

    The paper's diagnosis: "AI systems have cleared one celebrated benchmark after another... Yet by the metric that ultimately matters, economic output, the broader impact has remained surprisingly muted." ALE is "intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP-relevant impact."

    Implications

  • For developers: Long-horizon, cross-tool, domain-knowledge integration is the bottleneck; agent architecture (planning, memory, tool use, error recovery) matters more than raw model capability.
  • For enterprises: Benchmark scores are reference only — run POCs on real workflows; prioritize long-horizon stability over single-point success rates.
  • For researchers: ALE is a living benchmark; 13 of 55 subfields are entirely uncovered by existing benchmarks (Table 1 in the paper). Law, finance, and life sciences may be the next frontiers.
  • Limitations

    1. No physical-world tasks (robots, physical manufacturing operations). 2. Objective grading remains hard for some domains (creative writing, design aesthetics). 3. Static sandbox — real environments change (software updates, rule changes). 4. No safety-boundary evaluation.

    Conclusion

    Claude Fable 5's 80% on SWE-Bench vs. near-0% on ALE doesn't mean the model got worse — it means ALE measures completing whole jobs, not fixing single bugs. The next AI milestone isn't +5 points on a benchmark; it's moving full pass rates from 0% to a meaningful number.

    References

  • Sun, Yiyou et al. "Agents' Last Exam." arXiv:2606.05405 (2026). UC Berkeley RDI.
  • SWE-Bench Pro: https://www.swe-bench.com/
  • OSWorld: https://os-world.github.io/
  • Terminal-Bench: https://terminal-bench.com/
  • Berkeley RDI "How We Broke Top AI Agent Benchmarks" (April 2026)
  • Anthropic Claude Fable 5 Launch (June 9, 2026)
  • Scale AI SEAL Leaderboard: https://scale.com/leaderboard

Tags

#ai-benchmarks#agents-last-exam#ai-agents#swe-bench#uc-berkeley#agent-evaluation#real-world-workflows

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981495