What OpenAI shipped on August 13 wasn't just a new model card — it was a construction manual for running agents affordably. The GPT-5.6 family's selling point comes down to four words: price-performance.
The Cost Math
The most striking comparison: on BrowseComp (a fact-searching benchmark), GPT-5.5 (Extra High) scored 84.36% three months ago at a cost of $33.27. At launch, GPT-5.6 Luna (Extra High) scored 84.04% — for just $1.33.
Hypha's engineering lead put it even more bluntly: Luna retained 98% of GPT-5.5's extraction accuracy at one-eighteenth the cost.
More important is "architectural savings." GPT-5.6 was trained end-to-end for three things:
- Cross-turn reasoning persistence (keeping reasoning across turns)
- Native compaction of long conversations
- Programmatic tool calls (moving filtering, aggregation, and orchestration out of the model's context)
The Limits of Cheap
Cheapness has conditions. OpenAI repeatedly emphasizes tuning down reasoning effort: GPT-5.6 Sol at the low setting outperforms GPT-5.5 at high (same harness). The smaller Luna/Terra models suit high-throughput, low-latency, repetitive steps in agent pipelines — for example, legal-tech workflows that first extract from handwritten memos, then feed into frontier analysis. The prompt cache TTL has also been extended to at least 30 minutes with configurable deterministic breakpoints, helping a batch of startups cut uncached input by 28%.
Cheap Isn't the Endgame — It's the Turning Point
What this guide is really saying is that the economics of agents have changed. Tasks that once demanded a frontier model at every step can now match or beat results with a small model + tuned reasoning effort + architectural choices. When "expensive" is no longer a hard constraint for building agents, the bet shifts to who can design the most efficient harness — cost has become the core competitive capability of the new generation of AI coding.