Claude Fable 5.1 Deep Dive: Long-Horizon Agentic AI as the Main Battleground
*Research date: 2026-09-02. Subject: Anthropic Claude Fable 5.1 / Mythos 5.1 (released 2026-09-01). Primary source: Anthropic's official benchmark comparison table (7 benchmarks × 4 models, with 4 footnotes). Corroborating sources: Laude Institute / Stanford Terminal-Bench-Science 0.1 leaderboard, Zapier AutomationBench, Cursor CursorBench 3.2, Artificial Analysis (GDPval-AA v2, AutomationBench-AA), OSWorld 2.0 paper and Snorkel leaderboard, CAIS/Scale AI HLE.*
Headline numbers:
- 52.6% on Terminal-Bench-Science 0.1 — 2.13× over Fable 5
- 60.9% on HLE without tools — first public model above 60%
- 1853 GDPval-AA v2 Elo — the only loss (Opus 5: 1861)
- −75% cache-read pricing ($1.00 → $0.25 per million tokens)
- −45% estimated cost for highly agentic workloads
- 12 weeks — flagship lifespan of Fable 5 (May 20 → Sep 1)
- Fable 5.1 — publicly available; API ID
claude-fable-5-1; on AWS, Google Cloud, Azure. - Mythos 5.1 — same base model with looser guardrails; restricted to vetted cybersecurity and life-sciences organizations via verification programs; currently US-only, with Anthropic coordinating with the US government on expansion. Anthropic calls it its strongest cybersecurity model to date.
- Terminal-Bench-Science 0.1 (70 tasks from working scientists' pipelines, 7.6% acceptance rate): Fable 5.1's 52.6% smashes the previous public record of 30.0% (Opus 5) — the first time any model clears one-third, let alone half. Caveats: vendor harness (Laude's leaderboard scores Fable 5 at 21.4% vs Anthropic's 24.7% — a 3.3pp harness effect), and biology tasks were substituted to Opus 5 despite life sciences being the largest domain (19 tasks), so 52.6% may be *underestimated*.
- Terminal-Bench 4.0: +18.5pp over GPT-5.6 Sol, +13.8pp over Fable 5. Note v4.0 recalibrated quotas and is not comparable to earlier versions. A third-party-reported Mythos 5.1 score of 60.9% exactly matches Fable 5.1's HLE figure — flagged as possibly a data-scraping error.
- GDPval-AA v2 (OpenAI task set, Artificial Analysis agentic edition, 220 tasks, blind LLM-judged Elo, human baseline = 1000): The −8 Elo gap to Opus 5 equals a 51.2% pairwise win rate — statistically indistinguishable. Anthropic itself recommends starting with Opus 5 and moving to Fable 5.1 for stronger reasoning and long-horizon agentic work. Additional caveats: LLM judges carry known bias risk, and Opus 5's own Elo varied by 117 points across effort tiers in Artificial Analysis' August board.
- OSWorld 2.0 (108 end-to-end workflows, 31 apps, ~318 tool calls per task vs human ~1.6 hours): The sharpest derived metric is the strict/partial conversion rate: Fable 5.1 at 53.5% (vs Fable 5's 49.5%, Opus 4.8's 37.6%, GPT-5.5's 27.4%) — it doesn't just do more, it *finishes*. But 41.7% strict means 58.3% of long-horizon computer tasks still fail.
- Humanity's Last Exam: 60.9% without tools is the first publicly accessible model above 60% (historical: GPT-4o 2.7% in Jan 2025; Gemini 3.1 Pro 46.4%; prior best public Claude Opus 4.7 at 46.9%). Crucially, tool gains are nearly flat across all four models (+4.1 to +4.7pp), meaning the lead comes from raw reasoning, not tool use. Remaining caveats: with tools (65.0%) it only reaches the lower bound of cross-domain experts; and a 2025 FutureHouse review found 29% of HLE chemistry/biology answers contradicted literature.
- AutomationBench (Zapier; 47 simulated SaaS apps, ~500 API endpoints, ~12,000 assertions): 17.1% → 31.4% (+14.3pp, 1.84×) is the second-largest generational jump. But cross-vendor comparison largely fails here: private held-out task sets, different harnesses (Artificial Analysis' AutomationBench-AA gave Fable 5 a 48.6% — same model, 3× different by harness). Notably, every model tested by AutomationBench-AA violated business guardrails; Anthropic reported no guardrail data.
- CursorBench 3.2.0: The 73.4% "first place" carries the largest discount — only +2.9pp over Fable 5, inside Anthropic's own ±3.5pp error bar. The cost column tells the real story: Fable 5 (Max) cost $17.32/task (103,525 tokens, 72 steps) vs GPT-5.6 Sol's $5.69 (28,320 tokens, 48 steps) — 3× cost for +3.3pp. Fable 5.1's per-task cost is estimated at roughly $9.5–13 after the cache discount: still well above GPT-5.6 Sol.
One-Sentence Verdict
Fable 5.1 is not a minor patch. It is the first time Anthropic places "the ability to complete an entire task autonomously" at the center of its product pitch — first place on six of seven agentic benchmarks, with the single loss (GDPval-AA v2) within sampling error. The release is anchored commercially not by scores but by the 75% cache-read price cut (cutting long agentic-task bills by up to 45%) and enterprise frontier safeguards (EFS) that remove the mandatory 30-day data retention tradeoff, answering OpenAI's zero-retention positioning from three weeks earlier.
Timeline: Three Generations in Three Months
| Date | Event | Positioning | |---|---|---| | ~2026-05-20 | Fable 5 released | First "Fable" tier, above Opus | | 2026-07-24 | Opus 5 released (1M context default, adaptive thinking); GDPval-AA v2 at 1861 | 65 days after Fable 5 | | 2026-09-01 | Fable 5.1 + Mythos 5.1 released | First on 6 of 7 benchmarks, only 39 days after Opus 5 |
A flagship tier being replaced in three months compresses enterprise POC windows — model selection is now a quarterly, not annual, decision.
Twin Models: One Core, Two Guardrail Levels
Pricing: Input/Output Unchanged, Cache-Read Cut 75%
| Item | Fable 5 | Fable 5.1 | Change | |---|---|---|---| | Input (per M tokens) | $10.00 | $10.00 | — | | Output (per M tokens) | $50.00 | $50.00 | — | | Cache read | $1.00 | $0.25 | −75% | | 5-min / 1-hr cache write | — | $12.50 / $20.00 | — | | Context | 1M | 1M (adaptive thinking always on) | — |
Cache read falls to 0.025× of input price (vs 0.1× for other Claude models). Since every agent-loop iteration re-reads system prompts, history, and tool results, cache reads dominate long-task token bills. Official figures: ~25% total cost reduction for typical workloads, up to ~45% for highly agentic ones. Context: per payment platform Ramp, Fable 5 accounted for only ~11% of enterprise AI tool spend two months post-launch — cheaper Opus 5 out-earned it. Fable's bottleneck was never capability; it was the bill.
The Seven Benchmarks
| Domain | Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |---|---|---|---|---|---| | Agentic science | Terminal-Bench-Science 0.1 | 52.6% | 24.7% | – | 22.4% | | Agentic coding | Terminal-Bench 4.0 | 55.8% | 42.0% | 52.3% | 37.3% | | Knowledge work | GDPval-AA v2 (Elo) | 1853 | 1747 | 1861 | 1738 | | Computer use | OSWorld 2.0 (partial) | 77.9% | 72.9% | 70.6% | 62.6% | | Computer use | OSWorld 2.0 (strict) | 41.7% | 36.1% | – | – | | Reasoning | HLE (no tools) | 60.9% | 53.7% | 50.2% | 47.4% | | Reasoning | HLE (with tools) | 65.0% | 57.6% | 54.9% | 52.1% | | Business process | AutomationBench | 31.4% | 17.1% | – | – | | IDE coding | CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Anthropic's four footnotes (key to reading the table): 1. Tasks blocked by guardrails were substituted: cybersecurity tasks to Claude Opus 4.8, biology tasks to Opus 5 — likely *lowering* Fable 5.1/5 scores. 2. Terminal-Bench, AutomationBench, and CursorBench used Anthropic's own harness; cross-vendor numbers may not match other reports. 3. Fable 5.1 tested at max effort + adaptive thinking + 1M context; other models at their highest available settings. 4. Sampling n=100 (unless noted); margin of error ±3.5pp.
Key Findings by Benchmark
Bottom Line
Fable 5.1 converts long-horizon autonomy from a bonus feature into the core product claim, with genuine breakthroughs in autonomous science and open-model reasoning. But the honest reading: GDPval-AA is a tie, not a loss; CursorBench's win is within noise; OSWorld strict completion still leaves the majority of tasks unfinished; and economics — even after the 75% cache cut — remain the model's structural weakness against cheaper competitors.