Anthropic launched Claude Opus 5 on July 24 across all platforms. The core narrative in one sentence: pricing is unchanged from Opus 4.8, but the model runs within 0.5% of Fable 5's peak score on CursorBench 3.2, and on OSWorld 2.0 it surpasses Fable 5's best result at roughly one-third the cost.
What makes this worth a closer look is that it solidifies the "tiered pricing" of Anthropic's July model lineup. Opus 5 is priced at $5 input / $25 output per million tokens; Fable 5, according to multiple third-party platforms, sits at the $10 / $50 tier. Opus 5 hit SOTA on all four benchmark cards—CursorBench 3.2 max effort, Frontier-Bench v0.1, Zapier AutomationBench, and OSWorld 2.0. On Frontier-Bench v0.1 it delivered more than twice the performance of Opus 4.8 at lower per-task cost. On ARC-AGI 3 it scored roughly three times the second-place model.
More notable are several non-coding data points: a 22% improvement over Opus 4.7 on internal agentic coding evaluations; financial modeling accuracy up 9 percentage points with one-third fewer turns and 60% less time; and legal agent tasks using 26% fewer tokens than Opus 4.8 max reasoning. These aren't vague "stronger than Opus 4.8" claims—the differences can be converted into tool call counts and time.
But several boundaries need to be spelled out. First, Opus 5's cybersecurity classifier intervenes 85% less than Fable 5's, yet exploitation capability still falls far short of Mythos 5—vulnerability hunting is only permitted at the source-code level; binary scanning and penetration testing remain blocked by default. Second, the overall misalignment behavior score is 2.3, the lowest among recent generations; this widens the window for safely delegating longer task chains, but the official context window figure hasn't been disclosed. Third, the 22% agentic coding gain over Opus 4.7 comes from an internal evaluation whose specific design wasn't published; if you want a SWE-bench-style public comparison, you'll need to wait for external leaderboards.
Companion feature updates worth noting: the Claude API now offers "mid-conversation tool switching" (without invalidating the prompt cache) and "automatic fallback" (auto-routing to other models when blocked by a classifier). Both are details that reduce incident rates for teams running agents in production.
My own take: Opus 5 is the "underrated one" in Anthropic's July lineup. Fable 5's halo is too bright, and Opus 4.8 has had attention stolen by Mythos 5. But any engineering team that has run CursorBench / OSWorld / Frontier-Bench will rethink their budget once they see this pricing curve. Opus 5 is a key rung in the "model gradient convergence" of 2026 H2—for the first time, Anthropic simultaneously offers Fable 5 / Mythos 5's ceiling and Opus 5's affordable floor.