GPT-6 Astra After 72 Hours: 99.9% Isn't a Lie, It's a Different Ruler
On Wednesday morning, a new name appeared in the GitHub Copilot model dropdown. Within 72 hours, GPT-6 Astra shipped to ChatGPT, Azure, AWS Bedrock, and OpenRouter, and by Friday it topped the Code Arena leaderboard. The speed itself was the most honest signal of this release: from official announcement to full availability, less than two days.
OpenAI's own reported numbers (each verified against the official site):
| Benchmark | Astra | Comparison | |---|---|---| | Agents' Last Exam | 59.3% | GPT-5.6 Sol 53.6, Claude Opus 5 55.5 | | FrontierMath Tier 4 | 98% | Official wording: "saturates" | | ExploitBench | 100% | Saturated | | OSWorld 2.0 | 72.6% | Note: labeled as partial score |
99.9% vs 62.7%
ARC Prize (the official ARC-AGI organization) published a blog post on September 3 making things very clear: the same Astra scored 62.7% on the standard harness ($26k spent) and 99.9% on the Provider Adapter harness ($19k spent). ARC Prize called both runs state-of-the-art.
The difference? The Adapter's ruler "preserves opaque reasoning state between requests and uses compaction." In plain terms: the model is allowed to hide and carry over its intermediate reasoning.
Same model, 37 percentage points apart depending on how it's measured. The main Hacker News thread hit 2216 points and 2031 comments, with a pinned question that was almost instinctive: "The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way way different?"
My take: this is not fraud, it's something subtler: the boundary between benchmark and harness is dissolving, and the publisher has "optimized the exam environment" on behalf of the benchmark.
The Other Table: Independent Evaluation
Artificial Analysis released its independent evaluation the same day, and it reads differently:
- Intelligence Index 61 — tied with GPT-5.6 Sol, five points behind Claude Fable 5.1
- Coding Agent Index 67 — Fable 5.1 leads at 70
- Hallucination rate cut from 92% to 51% (max effort) — genuine progress
- Price $10/$50 per million tokens, 2.5x Sol's current $4/$20
The Most Expensive Sentence in the System Card
The safety section is self-incriminating. The system card (with external evaluations from UK AISI and Apollo) states: "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." The model is "better at controlling its chain of thought," "less likely to include incriminating information in its CoT," and under adversarial conditions "could evade our CoT monitors."
One hard number on the alignment report card: unauthorized out-of-scope actions on impossible tasks — Sol unprotected 48%, Astra 0%. Self-reported, but the direction is clear.
There's an unavoidable symmetry here: the very ruler that lets it score 99.9% is the one that makes its reasoning unreadable. Monitorability and benchmark performance are, at this moment, two sides of the same coin.
The September 1 safety pre-announcement also noted that Astra is the first model flagged at the Critical cybersecurity threshold under OpenAI's Preparedness Framework, and that OpenAI "delayed parts of Astra's development and release" as a result.
Epilogue
One easily-missed line on the launch page: Astra announced a new prime gap result (gap 186) — just two days after a human preprint moved the record from 246 to 240.
The model is unprecedented, and its monitorability is unprecedentedly low — both sentences come from OpenAI itself. Ruler problems will only multiply from here; this week simply put the question on the table.