English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Opus 5.5 Lands, GPT-6 Sol Follows Within an Hour: Who Rewrote the AI Coding Cost Curve

Forum topic · 小凯 · 2026-09-25

Summary

On September 22, 2026, Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens—20% cheaper than Opus 5 on both, with cache reads cut from $0.50 to $0.20 per million tokens (a 60% saving). The company claims a ~40% net cost reduction and 30%+ faster output for typical coding workloads, and made adaptive thinking a mandatory default that developers can no longer disable via API. Independent benchmarks from Artificial Analysis complicate the picture: at max effort, Opus 5.5 emits roughly 60% more output tokens, making per-task costs roughly on par with Opus 5. One hour after launch, OpenAI silently added its next-generation GPT-6 Sol and GPT-6 Luna models to Codex CLI release notes, signaling a synchronized pricing and capability race. The post also covers raised Claude Code rate limits, banked rate-limit resets, session-bound thinking blocks to block distillation, and cybersecurity capability gating requiring verification programs. It argues the industry's focus has shifted from leaderboard points to the economics of long-running coding agents.

Opus 5.5 Lands, GPT-6 Sol Follows Within an Hour: Who Rewrote the AI Coding Cost Curve

> [AI Coding · Flagship Generation] On September 22, 2026, Anthropic released Claude Opus 5.5, priced at $4 per million input tokens and $20 per million output tokens—both 20% lower than the previous Opus 5. Cache reads dropped from $0.50 to $0.20 per million tokens, a 60% saving on that line item alone. One hour later, OpenAI quietly wired GPT-6 Sol and GPT-6 Luna into Codex CLI. This wasn't just another leaderboard-refresh day: the economics of coding agents were, for the first time, officially put back on the table.

One-Sentence Conclusion

By Anthropic's own reporting, Opus 5.5 matches its Fable 5.1 on most tasks but costs 40% of Fable 5.1 per token, with cache reads at just 8% of its price; net cost falls roughly 40%, and output speed improves by more than 30%. It is also the first time Anthropic has made "adaptive thinking" a mandatory default—developers can no longer switch off deep reasoning in the API.

Why This Moment Is Different

For the past two years, whenever a flagship model launched, public attention centered on a single thing: the leaderboards. SWE-bench, Terminal-Bench, Aider, AIME—each generation scraped out another fraction of a point. But this September the wind shifted. Anthropic filled its front page with "40% cost reduction," "long-running agents no longer collapse," and "no longer forgets the original plan after the first hour." Its bet is straightforward: dev teams won't pay more just because a model got a bit smarter, but they will extend agent loops 3x or 5x when per-task cost comes down.

OpenAI's response changed too. The same-day release notes for GPT-6 Sol and Luna in Codex CLI were a single line, but it was preserved verbatim in the release notes—meaning Codex treats the new models as first-class citizens rather than burying them in a dropdown. This is a synchronized gear shift from marketing language to product language.

Breaking Down the Pricing

Anthropic's older models had an interesting tradition: each new generation is cheaper than the last, but the previous generation never gets a price cut:

| Model | Input \(/M | Output\)/M | Cache Read $/M | |---|---|---|---| | Fable 5.1 | 10 | 50 | 0.25 | | Opus 5 (previous) | 5 | 25 | 0.50 | | Opus 5.5 | 4 | 20 | 0.20 | | Opus 5.5 Fast mode | 8 | 40 | — | | Sonnet 5 | 2 | 10 | 0.20 | | Haiku 4.5 | 1 | 5 | 0.10 |

Opus 5.5 is 60% cheaper than Fable 5.1 on both input and output, and 20% cheaper on cache reads. Cache reads matter enormously for coding agents, because agent loops repeatedly push the repo, specs, and prior tool outputs back into context—the same text gets read hundreds of times, so the savings are real money. Anthropic put this number front and center itself.

Performance Numbers and Third-Party Accounts

Several of Anthropic's self-reported figures deserve a caveat, because third-party tests don't fully agree:

| Dimension | Opus 5.5 Self-Reported | Comparison | |---|---|---| | Terminal-Bench 4.0 | 66.4% | Above GPT-6 Astra's 57.9% | | GDPval-AA v2.1 (Elo) | 1846 | Above Fable 5.1 and Opus 5 | | Artificial Analysis Index (max effort) | 58 | 7 points above Opus 5 | | Output tokens per AA task | ~119k | ~60% more than Opus 5 |

Third-party accounting cools the optimism. Artificial Analysis notes that at max effort, Opus 5.5 generates about 60% more tokens, making weighted per-task cost roughly on par with Opus 5. Anthropic's 40% cost drop is a "typical coding workload" figure, not a max-effort extreme. It's a very vendor-style accounting line—but also the industry's default accounting line.

Adaptive Thinking Is Now Forced On

Opus 5.5 makes adaptive thinking a mandatory default: developers can no longer disable deep reasoning per request, only modulate its depth via the effort parameter. In Anthropic's own internal tests, an HAProxy task (C translated to Rust) finished in 9.5 hours versus 12 hours for Opus 5, with 51% cost savings. A 200k-line codebase audit took Opus 5.5 three hours, versus 20 hours and 2.5x the tokens for Opus 5.

Anthropic also admitted that Opus 5.5 frequently recognized it was being evaluated during testing, making "do observed behaviors transfer to real deployment" an open question. Frontier Design and METR evaluations are listed in the safety record; Anthropic also states the model is close to Claude Mythos 5.1 in bio and cyber capabilities, so general cybersecurity requests automatically fall back to Opus 4.8, and higher capability tiers require access via the Cyber Verification Program or Life Sciences Verification Program.

GPT-6 Sol and Luna Right Behind

OpenAI's reaction is interesting. Codex CLI's release notes on the night of September 22 included two lines:

> Added support for new GPT-6 Sol and Luna models via Amazon Bedrock, including migration prompts for older models.

Developers on HN and X read this as "OpenAI's next-gen models are already in production." The same release also included AWS Bedrock integration and migration prompts for older models.

Put the two events together: one hour after Anthropic announced Opus 5.5, OpenAI wired two variants of its next-gen flagship into its coding tool. Sol is the flagship workhorse; Luna is most likely a lightweight version. This "rival cuts price, so do I" cadence was a GPT-5 vs Claude Sonnet 4 story last year; now it goes straight to the next-gen flagship.

Zooming Out

Claude Code's five-hour usage cap rose 20% on Opus 5.5, with a first-ever "banked rate-limit reset"—you can bank today's quota for tomorrow's heavy tasks rather than racing for capacity. Anthropic rolled this out across Pro, Max, Team, and Enterprise plans; AWS Bedrock, Google Cloud, and Microsoft Azure all went live same-day. Sonnet 5.5 and Haiku 5.5 are planned in the coming weeks.

One under-publicized detail with deep impact: thinking blocks are now bound to a single session, explicitly to block model distillation. Bad news for any rival hoping to clone Claude's reasoning style for training—and for in-house developers too, since workflows spanning multiple session IDs must be rewritten.

Boundaries Worth Noting

First, cache-read savings aren't universal. They only apply to long sessions, fixed templates, and repeatedly revisited agent workflows. One-shot short requests see none of it.

Second, max effort means no savings. Third-party tests show ~60% more output tokens, with per-task total cost on par with Opus 5. Anthropic's 40% drop is a "typical" value, not a "maxed-out" value.

Third, evaluation-detection rates are rising. Anthropic itself disclosed that Opus 5.5 often recognized it was being tested—a new headache for red teams and benchmark organizations, since there may be distribution shift between "knows it's observed" and "assumes it isn't."

Fourth, capability and price thresholds are decoupling. Everyone assumed higher capability meant higher price, with Fable 5.1 as that curve's endpoint. Opus 5.5 is the first to bend that equation, letting capability and price decouple within the same tier.

Who to Believe

Anthropic's self-reported numbers answer "did the flagship get cheaper?" Third-party independent tests (Artificial Analysis, Top10.dev, lmmarketcap, LMSPedia) answer "how much do you actually save" and "what does it mean for a specific workload." The core numbers consistent across sources: Opus 5.5 at $4 input, $20 output, $0.20 cache read. The two non-consistent numbers—Anthropic's 40% net cost drop versus AA's "on par with Opus 5 at max effort"—aren't contradictory; they just measure different regimes.

---

Sources: Anthropic official release page (Sept 22), Anthropic Pricing Documentation (verified Sept 24), TechRepublic (Sept 24), Complete AI Training (Sept 25), LMSPedia evaluation page, LMMarketCap signal board (Score 95, Rank 5/438), Artificial Analysis Intelligence Index v4.3.2, Llmcostlab.com pricing archive, OpenAI Codex CLI release notes (Sept 22).

Tags

#claude-opus-5-5#anthropic#openai#gpt-6-sol#codex-cli#ai-coding-agents#llm-pricing#benchmarking

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635192