Codex's New Division of Labor: Sol as Foreman, Luna Max as Bounded Manual Laborer
Category: tip · AI coding / Agent Harness Time: 2026-08-02 10:47 (UTC+8) Sources: OpenAI official docs and pricing notes; community workflow repo; AYi's public hands-on test
A Community Workflow Goes Viral
On August 2, AYi shared a Codex usage pattern: the main thread runs GPT-5.6 Sol, responsible for breaking down tasks, making architectural judgments, and final review; clearly specified, verifiable implementation, bug fixes, test runs, and refactoring are delegated to GPT-5.6 Luna Max subagents. In other words, Sol acts as the foreman and Luna as the worker.
This is not an officially announced new product — it's a workflow the community assembled from two capabilities that already exist. OpenAI's official documentation confirms that Codex supports parallel subagents and custom agent configuration. Agent files live in ~/.codex/agents/ or a project's .codex/agents/, must define name, description, and developer_instructions, and can optionally specify model and model_reasoning_effort individually.
Why This Layering Makes Economic Sense
OpenAI's July 30 pricing announcement provides firmer ground: GPT-5.6 Luna API pricing dropped to $0.20 per million input tokens and $1.20 per million output tokens; Sol pricing is unchanged. Luna can still call tools and complete multi-step workflows, making it well suited to absorbing large volumes of well-bounded tasks.
The official coding example follows nearly the same structure: Sol first handles uncertainty and defines the plan, then lets Luna implement explicit changes, write tests, run tests, and evaluate results. The community repository use-luna-subagents turns this idea into a reviewable configuration: showing diffs before installation, pinning the Luna model, preserving the parent session's sandbox and permissions, forbidding silent model substitution, and requiring the parent agent to verify results.
Don't Treat "Doubled Output" as a Benchmark Result
AYi's claim that "output per subscription directly doubles" is a personal workflow impression, not a published controlled experiment. Official documentation instead warns about three easily overlooked issues:
- Subagents increase token consumption
- Permissions and sandboxing are inherited from the parent agent by default
- Parallel writes with overlapping scopes create conflicts
- https://x.com/AYi_AInotes/status/2083867265179537565
- https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6
- https://developers.openai.com/codex/subagents
- https://github.com/aitransformationdirector/use-luna-subagents
A safer way to adopt this: first have Sol delegate only rollback-able, clearly bounded, independently testable tasks; Luna does not directly modify shared configs or critical interfaces; the main thread must review diffs, run tests, and then merge. Complex design, cross-module migrations, and security judgments remain with Sol.
The significance of this development is not that "cheap models suddenly equal expensive ones," but that the unit of cost in AI coding is shifting from "one conversation" to "a set of layered agents plus a verification process." Model routing and harness design are replacing pure model selection as the core of productivity.
Original post and evidence: