Superpowers v6 Deep Dive: Fable-Driven 36-Hour Autonomous R&D Delivers 50% Faster Builds and 60% Cost Cuts
> Author: Jesse Vincent (Prime Radiant) > Published: 2026-06-15 > Project: https://github.com/obra/superpowers > Evals: https://github.com/prime-radiant-inc/superpowers-evals > Original post: https://primeradiant.com/blog/2026/superpowers-6.html
In one line: Superpowers skipped from 5.2 straight to 6 — not because of feature accumulation, but because Fable ran 25 quantified experiments in 36 hours, fundamentally restructuring the cost profile of Subagent Driven Development.
1. From 5.2 to 6: A Release Plan Rewritten by Fable
Superpowers was set to ship as 5.2. The release had been delayed twice while the team added Pi/Antigravity/Kimi Code support, rewrote model-agnostic skills, improved Visual Brainstorming, and fixed a pile of bugs.
Then Anthropic released Fable.
Within days of getting access, founder Jesse Vincent did something highly characteristic: instead of having Fable write new features, he had Fable optimize Superpowers' own build loop. The results far exceeded expectations — not 15% token savings, but 50% wall-clock speedup and 60% token cost reduction.
That was enough to bump the version number from 5.2 to 6.
2. Superpowers' Core Workflow, Reviewed
Before dissecting v6, recall Superpowers' standard pipeline:
1. Brainstorming: visual brainstorming, outputting a design doc 2. Isolated workspaces: a separate branch/environment per task 3. Plan writing: a detailed implementation spec 4. TDD-driven implementation: strict red-green test-driven development 5. Code review: two gates — code quality review + spec compliance review 6. Merge on completion: branches merge only after passing review
This pipeline guarantees quality but imposes two costs: it is slow (layered reviews) and expensive (every step consumes heavy tokens).
Jesse's own words: "It's never made me happy that it's slow and expensive."
3. Three Core Breakthroughs in v6
Breakthrough 1: Coordinator–Reviewer Handoff Optimization — Pre-baked Review Packets
Problem: Review subagents often ran many git commands just to obtain the diff to review.
Fix: Convert "how to find the commits to review" from text instructions into a shell script that pre-generates a review packet — well-formatted diffs plus metadata.
Result: ~10% reduction in both token consumption and wall-clock time.
The elegance of this optimization: it doesn't make the reviewer "smarter" — it completes in advance the work the reviewer would have done (querying git, formatting diffs), letting the reviewer focus on judgment rather than preparation.
Breakthrough 2: Merging the Reviewer Architecture — From Two Gates to One
Problem: Superpowers previously had two independent review agents:
- Code Quality Reviewer: is the code quality up to standard?
- Spec Compliance Reviewer: does the implementation strictly match the spec?
- A unified two-stage review prompt rewrite
- Introduction of a highly instructive "unverifiable from diff" state
- Stricter review that is harder to bypass
- A complete autoresearch harness
- 25 closed experiments + 4 backlogged
- Pre-registered predictions for every hypothesis
- Full experiment log:
docs/experiments/2026-06-11-build-loop-autoresearch.md - Loop cost: ~$165 (~$650 at unsubsidized token prices)
- Turns rose from 92 to 138; output doubled
- Conclusion: thinking buys turn efficiency; capping it doesn't pay
- Even with code sections exempted, test content was cut -62%
- Fidelity held, but task structure caved in
- Tests + interfaces + structure already carried the full load
- The reviewer produces confident spec verdicts
- But silently redefines "spec" as "global constraints"
- 0/5 flagged a missing brief
- report reads
- cache health
- reviewer floor
- haiku fixers
- todo bookkeeping
- dispatch re-derivation
- Validates improvements across multiple harnesses (Claude, Codex, OpenCode, Cursor)
- Quantifies each change's impact on build speed and token cost
- Surfaced a "false negative" on Codex — the initial Codex eval wasn't sufficiently isolated from the host OS and had been testing 5.1.0 all along
- N=5 gate battery still pending: E27's optimal config needs validation on larger samples
- Reviewer variance is a quality risk: the terse contract's effectiveness depends on reviewer consistency
- Diff-only review risks missing context: reviewers have been observed silently redefining "spec"
- Codex isolation issues: correct eval environment configuration is a precondition, or the optimization gains may be illusory
- Project: https://github.com/obra/superpowers
- Evals repo: https://github.com/prime-radiant-inc/superpowers-evals
- Official blog: https://primeradiant.com/blog/2026/superpowers-6.html
- Key person: Jesse Vincent (Prime Radiant founder, Superpowers author)
- Key tool: Anthropic Fable (used for the autoresearch loop)
- Experiment log:
docs/experiments/2026-06-11-build-loop-autoresearch.md - Published: 2026-06-15
Jesse's intuition: merge the two reviewers into one — could that save 15%?
Fable's action: Jesse posted the idea to internal Slack before bed. Fable independently reached the same conclusion overnight, tested it, and by morning the results were sitting in the evals.
Result: an additional ~15% savings in wall-clock time and tokens.
Key technical details:
Breakthrough 3: Fable's Autonomous R&D Loop — 25 Experiments in 36 Hours
This is the soul of v6. Jesse's instruction to Fable:
> /goal once this is done, run an autoresearch loop to improve cost-efficiency of the superpowers build loop. test with opus as the coordinator. make an hypothesis log. run experiments. run at least 25 experiments.
What Fable delivered:
4. Key Findings from the 25 Experiments
🏆 The ship candidate (E27) — Optimal Configuration
Combo: opus controller + elicited plan + conditional haiku implementers + terse reviewer contract + narration recipe + final-review tier pin
Cost: fractals $6.24/$6.60 (vs. baseline combo configs at $11.67–14.84)
Quality: planted-defect gates 2/3, with the sole failure attributed to reviewer variance + judge strictness; the terse contract was explicitly exonerated
📈 Quantified Wins
| Optimization | Effect | |---|---| | Terse reviewer contract | -41% reviewer output, verdict completeness preserved | | Narration recipe | -54% output, zero variance | | Conditional implementer tiering | ~$0.5–1/run saved; correctly refuses to let haiku handle prose plans | | Fixture-realism on svelte | -24% scope-matched |
💀 Intuitions Proven "Dead"
These findings are extremely valuable — directions proven ineffective by experiment, saving the team future investment:
1. Capping controller thinking → counterproductive
2. Plan word budgets → gutted test content
3. Sonnet plan generation → task structure collapse
4. Implementation bodies in plans → near-zero marginal utility
⚠️ A Risk Finding Worth Remembering
Problems when the reviewer gets only a diff package:
Jesse's take: "Same failure family as the haiku-reviewer advocacy."
This is a warning for any team automating code review: the context boundary given to a reviewer must be precisely defined, or the reviewer will confidently reach wrong conclusions.
✅ Six Directions Already Optimal
These were proven already optimal, recorded to prevent re-paying for them:
🔧 Methodological Lessons
Fable caught 3 measurement bugs in its own loop: 1. grep counted template echoes as self-review catches 2. the harness never inlined diffs 3. the scorer regex missed a newline
One verdict was retracted and re-measured: -74% corrected to an honest -41%.
This demonstrates a key discipline of autoresearch: manual inspection isn't optional — it's mandatory.
5. The Eval System: From "Feels Faster" to "Measurable"
Without the eval system, the v6 release would have been a mess of impressions.
Value of the Superpowers evals:
GitHub: https://github.com/prime-radiant-inc/superpowers-evals
The discovery process itself is instructive: if your eval environment isn't properly isolated, you may be optimizing a stale version.
6. General Design Principles from Superpowers v6
1. Handoffs Are Cost
The Coordinator→Reviewer handoff is a natural cost point. v6's approach isn't "make the handoff faster" but "reduce the reviewer's information-preparation burden" — pre-baked packets make the review ready to consume on arrival.
2. Merging Review Gates Requires Prompt Engineering
Merging two reviewers isn't simple subtraction. The keys were the unified two-stage review prompt rewrite and the "unverifiable from diff" state for edge cases.
3. Autoresearch Is the Right Way to Optimize Cost
36 hours, $165, 25 experiments — this isn't "let the AI try stuff," it's the scientific method: pre-registered hypotheses + systematic measurement + recorded failures.
4. Proving What Doesn't Work Matters as Much as Proving What Does
In v6's autoresearch log, "dead ideas" and "wins" are treated with equal weight. The team never has to waste another day on capping controller thinking or plan word budgets.
7. Limitations and Risks
8. Closing: The Next Benchmark for AI Coding Workflows
Superpowers v6 isn't another "support more models" feature update. It is the first AI coding workflow framework systematically optimized by a Fable-class autonomous research loop.
The 50% speedup and 60% cost reduction weren't bought by sacrificing quality — planted-defect gates still pass 2/3, with the sole failure attributed to reviewer variance rather than the prompts themselves.
As Jesse Vincent put it at the end of his post: "We're very proud of the improvements that we (and our robot buddies) have made."
"Robot buddies" is the right phrase. Superpowers v6 is the product of human engineers collaborating with an AI autonomous research loop. The 36-hour autoresearch run doesn't replace human judgment — it frees humans from 25 rounds of trial and error, letting them focus on judging which findings are worth shipping.
For every team building AI coding workflows, Superpowers v6's experiment log (docs/experiments/2026-06-11-build-loop-autoresearch.md) is a free experience pack — recording what works, what doesn't, and why.
References