English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Claude Fable 5.1 Deep Dive: Anthropic Makes Long-Horizon Agentic AI Its Main Battleground

Forum topic · ✨步子哥 · 2026-09-01

Summary

A detailed analysis of Anthropic's Claude Fable 5.1 and Mythos 5.1 (released 2026-09-01), based on Anthropic's official benchmark tables plus third-party leaderboards. Fable 5.1 ranks first on six of seven agentic benchmarks, most notably jumping from 24.7% to 52.6% on Terminal-Bench-Science 0.1 and becoming the first publicly available model to exceed 60% on Humanity's Last Exam without tools (60.9%). Its only loss is GDPval-AA v2 (1853 Elo vs Opus 5's 1861), a statistically insignificant gap. The article also highlights a 75% cut in cache-read pricing ($1.00 to $0.25 per million tokens), cutting total costs roughly 25% and up to 45% for highly agentic workloads, and notes the twin-model strategy: Fable 5.1 is publicly available while the looser-guardrailed Mythos 5.1 serves vetted cybersecurity and life-sciences organizations. Critical caveats are examined throughout: vendor-run harnesses, sampling error of ±3.5pp at n=100, LLM-judged Elo rankings, guardrail substitutions in scoring, and OSWorld 2.0's strict completion rate still at 41.7%.

Claude Fable 5.1 Deep Dive: Long-Horizon Agentic AI as the Main Battleground

*Research date: 2026-09-02. Subject: Anthropic Claude Fable 5.1 / Mythos 5.1 (released 2026-09-01). Primary source: Anthropic's official benchmark comparison table (7 benchmarks × 4 models, with 4 footnotes). Corroborating sources: Laude Institute / Stanford Terminal-Bench-Science 0.1 leaderboard, Zapier AutomationBench, Cursor CursorBench 3.2, Artificial Analysis (GDPval-AA v2, AutomationBench-AA), OSWorld 2.0 paper and Snorkel leaderboard, CAIS/Scale AI HLE.*

Headline numbers:

  • 52.6% on Terminal-Bench-Science 0.1 — 2.13× over Fable 5
  • 60.9% on HLE without tools — first public model above 60%
  • 1853 GDPval-AA v2 Elo — the only loss (Opus 5: 1861)
  • −75% cache-read pricing ($1.00 → $0.25 per million tokens)
  • −45% estimated cost for highly agentic workloads
  • 12 weeks — flagship lifespan of Fable 5 (May 20 → Sep 1)
  • One-Sentence Verdict

    Fable 5.1 is not a minor patch. It is the first time Anthropic places "the ability to complete an entire task autonomously" at the center of its product pitch — first place on six of seven agentic benchmarks, with the single loss (GDPval-AA v2) within sampling error. The release is anchored commercially not by scores but by the 75% cache-read price cut (cutting long agentic-task bills by up to 45%) and enterprise frontier safeguards (EFS) that remove the mandatory 30-day data retention tradeoff, answering OpenAI's zero-retention positioning from three weeks earlier.

    Timeline: Three Generations in Three Months

    | Date | Event | Positioning | |---|---|---| | ~2026-05-20 | Fable 5 released | First "Fable" tier, above Opus | | 2026-07-24 | Opus 5 released (1M context default, adaptive thinking); GDPval-AA v2 at 1861 | 65 days after Fable 5 | | 2026-09-01 | Fable 5.1 + Mythos 5.1 released | First on 6 of 7 benchmarks, only 39 days after Opus 5 |

    A flagship tier being replaced in three months compresses enterprise POC windows — model selection is now a quarterly, not annual, decision.

    Twin Models: One Core, Two Guardrail Levels

  • Fable 5.1 — publicly available; API ID claude-fable-5-1; on AWS, Google Cloud, Azure.
  • Mythos 5.1 — same base model with looser guardrails; restricted to vetted cybersecurity and life-sciences organizations via verification programs; currently US-only, with Anthropic coordinating with the US government on expansion. Anthropic calls it its strongest cybersecurity model to date.
  • Pricing: Input/Output Unchanged, Cache-Read Cut 75%

    | Item | Fable 5 | Fable 5.1 | Change | |---|---|---|---| | Input (per M tokens) | $10.00 | $10.00 | — | | Output (per M tokens) | $50.00 | $50.00 | — | | Cache read | $1.00 | $0.25 | −75% | | 5-min / 1-hr cache write | — | $12.50 / $20.00 | — | | Context | 1M | 1M (adaptive thinking always on) | — |

    Cache read falls to 0.025× of input price (vs 0.1× for other Claude models). Since every agent-loop iteration re-reads system prompts, history, and tool results, cache reads dominate long-task token bills. Official figures: ~25% total cost reduction for typical workloads, up to ~45% for highly agentic ones. Context: per payment platform Ramp, Fable 5 accounted for only ~11% of enterprise AI tool spend two months post-launch — cheaper Opus 5 out-earned it. Fable's bottleneck was never capability; it was the bill.

    The Seven Benchmarks

    | Domain | Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |---|---|---|---|---|---| | Agentic science | Terminal-Bench-Science 0.1 | 52.6% | 24.7% | – | 22.4% | | Agentic coding | Terminal-Bench 4.0 | 55.8% | 42.0% | 52.3% | 37.3% | | Knowledge work | GDPval-AA v2 (Elo) | 1853 | 1747 | 1861 | 1738 | | Computer use | OSWorld 2.0 (partial) | 77.9% | 72.9% | 70.6% | 62.6% | | Computer use | OSWorld 2.0 (strict) | 41.7% | 36.1% | – | – | | Reasoning | HLE (no tools) | 60.9% | 53.7% | 50.2% | 47.4% | | Reasoning | HLE (with tools) | 65.0% | 57.6% | 54.9% | 52.1% | | Business process | AutomationBench | 31.4% | 17.1% | – | – | | IDE coding | CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |

    Anthropic's four footnotes (key to reading the table): 1. Tasks blocked by guardrails were substituted: cybersecurity tasks to Claude Opus 4.8, biology tasks to Opus 5 — likely *lowering* Fable 5.1/5 scores. 2. Terminal-Bench, AutomationBench, and CursorBench used Anthropic's own harness; cross-vendor numbers may not match other reports. 3. Fable 5.1 tested at max effort + adaptive thinking + 1M context; other models at their highest available settings. 4. Sampling n=100 (unless noted); margin of error ±3.5pp.

    Key Findings by Benchmark

  • Terminal-Bench-Science 0.1 (70 tasks from working scientists' pipelines, 7.6% acceptance rate): Fable 5.1's 52.6% smashes the previous public record of 30.0% (Opus 5) — the first time any model clears one-third, let alone half. Caveats: vendor harness (Laude's leaderboard scores Fable 5 at 21.4% vs Anthropic's 24.7% — a 3.3pp harness effect), and biology tasks were substituted to Opus 5 despite life sciences being the largest domain (19 tasks), so 52.6% may be *underestimated*.
  • Terminal-Bench 4.0: +18.5pp over GPT-5.6 Sol, +13.8pp over Fable 5. Note v4.0 recalibrated quotas and is not comparable to earlier versions. A third-party-reported Mythos 5.1 score of 60.9% exactly matches Fable 5.1's HLE figure — flagged as possibly a data-scraping error.
  • GDPval-AA v2 (OpenAI task set, Artificial Analysis agentic edition, 220 tasks, blind LLM-judged Elo, human baseline = 1000): The −8 Elo gap to Opus 5 equals a 51.2% pairwise win rate — statistically indistinguishable. Anthropic itself recommends starting with Opus 5 and moving to Fable 5.1 for stronger reasoning and long-horizon agentic work. Additional caveats: LLM judges carry known bias risk, and Opus 5's own Elo varied by 117 points across effort tiers in Artificial Analysis' August board.
  • OSWorld 2.0 (108 end-to-end workflows, 31 apps, ~318 tool calls per task vs human ~1.6 hours): The sharpest derived metric is the strict/partial conversion rate: Fable 5.1 at 53.5% (vs Fable 5's 49.5%, Opus 4.8's 37.6%, GPT-5.5's 27.4%) — it doesn't just do more, it *finishes*. But 41.7% strict means 58.3% of long-horizon computer tasks still fail.
  • Humanity's Last Exam: 60.9% without tools is the first publicly accessible model above 60% (historical: GPT-4o 2.7% in Jan 2025; Gemini 3.1 Pro 46.4%; prior best public Claude Opus 4.7 at 46.9%). Crucially, tool gains are nearly flat across all four models (+4.1 to +4.7pp), meaning the lead comes from raw reasoning, not tool use. Remaining caveats: with tools (65.0%) it only reaches the lower bound of cross-domain experts; and a 2025 FutureHouse review found 29% of HLE chemistry/biology answers contradicted literature.
  • AutomationBench (Zapier; 47 simulated SaaS apps, ~500 API endpoints, ~12,000 assertions): 17.1% → 31.4% (+14.3pp, 1.84×) is the second-largest generational jump. But cross-vendor comparison largely fails here: private held-out task sets, different harnesses (Artificial Analysis' AutomationBench-AA gave Fable 5 a 48.6% — same model, 3× different by harness). Notably, every model tested by AutomationBench-AA violated business guardrails; Anthropic reported no guardrail data.
  • CursorBench 3.2.0: The 73.4% "first place" carries the largest discount — only +2.9pp over Fable 5, inside Anthropic's own ±3.5pp error bar. The cost column tells the real story: Fable 5 (Max) cost $17.32/task (103,525 tokens, 72 steps) vs GPT-5.6 Sol's $5.69 (28,320 tokens, 48 steps) — 3× cost for +3.3pp. Fable 5.1's per-task cost is estimated at roughly $9.5–13 after the cache discount: still well above GPT-5.6 Sol.

Bottom Line

Fable 5.1 converts long-horizon autonomy from a bonus feature into the core product claim, with genuine breakthroughs in autonomous science and open-model reasoning. But the honest reading: GDPval-AA is a tie, not a loss; CursorBench's win is within noise; OSWorld strict completion still leaves the majority of tasks unfinished; and economics — even after the 75% cache cut — remain the model's structural weakness against cheaper competitors.

Tags

#anthropic#claude-fable-5-1#agentic-ai#benchmarks#terminal-bench-science#humanitys-last-exam#osworld#ai-pricing

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634383