English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Superpowers v6 Deep Dive: Fable-Driven 36-Hour Autonomous R&D Delivers 50% Faster Builds and 60% Cost Cuts

Forum topic · 小凯 · 2026-07-06

Summary

Superpowers, Jesse Vincent's subagent-driven development framework, jumped from v5.2 straight to v6 after an autonomous research loop run by Anthropic's Fable completed 25 quantitative experiments in 36 hours for roughly $165 in tokens. The result: a 50% reduction in wall-clock build time and a 60% reduction in token cost, without sacrificing quality. Key changes include pre-baked review packets that eliminate reviewer preparation overhead (~10% savings), merging the code-quality and spec-compliance reviewers into a single two-stage reviewer with a new 'unverifiable from diff' state (~15% savings), terse reviewer contracts (-41% reviewer output), narration recipes (-54% output with zero variance), and conditional implementer model tiering. Equally valuable are the disproven ideas: capping controller thinking backfired (turns rose from 92 to 138), plan word budgets gutted test content by 62%, and Sonnet-generated plans collapsed task structure. A critical risk finding: reviewers given only diff packages confidently issued verdicts while silently redefining 'spec' as 'global constraints', missing missing briefs 0/5 times. The framework also caught three measurement bugs in its own harness, underscoring that manual inspection is mandatory in autonomous research. Full experiment logs and evals are open-sourced on GitHub.

Superpowers v6 Deep Dive: Fable-Driven 36-Hour Autonomous R&D Delivers 50% Faster Builds and 60% Cost Cuts

> Author: Jesse Vincent (Prime Radiant) > Published: 2026-06-15 > Project: https://github.com/obra/superpowers > Evals: https://github.com/prime-radiant-inc/superpowers-evals > Original post: https://primeradiant.com/blog/2026/superpowers-6.html

In one line: Superpowers skipped from 5.2 straight to 6 — not because of feature accumulation, but because Fable ran 25 quantified experiments in 36 hours, fundamentally restructuring the cost profile of Subagent Driven Development.

1. From 5.2 to 6: A Release Plan Rewritten by Fable

Superpowers was set to ship as 5.2. The release had been delayed twice while the team added Pi/Antigravity/Kimi Code support, rewrote model-agnostic skills, improved Visual Brainstorming, and fixed a pile of bugs.

Then Anthropic released Fable.

Within days of getting access, founder Jesse Vincent did something highly characteristic: instead of having Fable write new features, he had Fable optimize Superpowers' own build loop. The results far exceeded expectations — not 15% token savings, but 50% wall-clock speedup and 60% token cost reduction.

That was enough to bump the version number from 5.2 to 6.

2. Superpowers' Core Workflow, Reviewed

Before dissecting v6, recall Superpowers' standard pipeline:

1. Brainstorming: visual brainstorming, outputting a design doc 2. Isolated workspaces: a separate branch/environment per task 3. Plan writing: a detailed implementation spec 4. TDD-driven implementation: strict red-green test-driven development 5. Code review: two gates — code quality review + spec compliance review 6. Merge on completion: branches merge only after passing review

This pipeline guarantees quality but imposes two costs: it is slow (layered reviews) and expensive (every step consumes heavy tokens).

Jesse's own words: "It's never made me happy that it's slow and expensive."

3. Three Core Breakthroughs in v6

Breakthrough 1: Coordinator–Reviewer Handoff Optimization — Pre-baked Review Packets

Problem: Review subagents often ran many git commands just to obtain the diff to review.

Fix: Convert "how to find the commits to review" from text instructions into a shell script that pre-generates a review packet — well-formatted diffs plus metadata.

Result: ~10% reduction in both token consumption and wall-clock time.

The elegance of this optimization: it doesn't make the reviewer "smarter" — it completes in advance the work the reviewer would have done (querying git, formatting diffs), letting the reviewer focus on judgment rather than preparation.

Breakthrough 2: Merging the Reviewer Architecture — From Two Gates to One

Problem: Superpowers previously had two independent review agents:

  • Code Quality Reviewer: is the code quality up to standard?
  • Spec Compliance Reviewer: does the implementation strictly match the spec?
  • Jesse's intuition: merge the two reviewers into one — could that save 15%?

    Fable's action: Jesse posted the idea to internal Slack before bed. Fable independently reached the same conclusion overnight, tested it, and by morning the results were sitting in the evals.

    Result: an additional ~15% savings in wall-clock time and tokens.

    Key technical details:

  • A unified two-stage review prompt rewrite
  • Introduction of a highly instructive "unverifiable from diff" state
  • Stricter review that is harder to bypass
  • Breakthrough 3: Fable's Autonomous R&D Loop — 25 Experiments in 36 Hours

    This is the soul of v6. Jesse's instruction to Fable:

    > /goal once this is done, run an autoresearch loop to improve cost-efficiency of the superpowers build loop. test with opus as the coordinator. make an hypothesis log. run experiments. run at least 25 experiments.

    What Fable delivered:

  • A complete autoresearch harness
  • 25 closed experiments + 4 backlogged
  • Pre-registered predictions for every hypothesis
  • Full experiment log: docs/experiments/2026-06-11-build-loop-autoresearch.md
  • Loop cost: ~$165 (~$650 at unsubsidized token prices)
  • 4. Key Findings from the 25 Experiments

    🏆 The ship candidate (E27) — Optimal Configuration

    Combo: opus controller + elicited plan + conditional haiku implementers + terse reviewer contract + narration recipe + final-review tier pin

    Cost: fractals $6.24/$6.60 (vs. baseline combo configs at $11.67–14.84)

    Quality: planted-defect gates 2/3, with the sole failure attributed to reviewer variance + judge strictness; the terse contract was explicitly exonerated

    📈 Quantified Wins

    | Optimization | Effect | |---|---| | Terse reviewer contract | -41% reviewer output, verdict completeness preserved | | Narration recipe | -54% output, zero variance | | Conditional implementer tiering | ~$0.5–1/run saved; correctly refuses to let haiku handle prose plans | | Fixture-realism on svelte | -24% scope-matched |

    💀 Intuitions Proven "Dead"

    These findings are extremely valuable — directions proven ineffective by experiment, saving the team future investment:

    1. Capping controller thinking → counterproductive

  • Turns rose from 92 to 138; output doubled
  • Conclusion: thinking buys turn efficiency; capping it doesn't pay
  • 2. Plan word budgets → gutted test content

  • Even with code sections exempted, test content was cut -62%
  • 3. Sonnet plan generation → task structure collapse

  • Fidelity held, but task structure caved in
  • 4. Implementation bodies in plans → near-zero marginal utility

  • Tests + interfaces + structure already carried the full load
  • ⚠️ A Risk Finding Worth Remembering

    Problems when the reviewer gets only a diff package:

  • The reviewer produces confident spec verdicts
  • But silently redefines "spec" as "global constraints"
  • 0/5 flagged a missing brief
  • Jesse's take: "Same failure family as the haiku-reviewer advocacy."

    This is a warning for any team automating code review: the context boundary given to a reviewer must be precisely defined, or the reviewer will confidently reach wrong conclusions.

    ✅ Six Directions Already Optimal

    These were proven already optimal, recorded to prevent re-paying for them:

  • report reads
  • cache health
  • reviewer floor
  • haiku fixers
  • todo bookkeeping
  • dispatch re-derivation
  • 🔧 Methodological Lessons

    Fable caught 3 measurement bugs in its own loop: 1. grep counted template echoes as self-review catches 2. the harness never inlined diffs 3. the scorer regex missed a newline

    One verdict was retracted and re-measured: -74% corrected to an honest -41%.

    This demonstrates a key discipline of autoresearch: manual inspection isn't optional — it's mandatory.

    5. The Eval System: From "Feels Faster" to "Measurable"

    Without the eval system, the v6 release would have been a mess of impressions.

    Value of the Superpowers evals:

  • Validates improvements across multiple harnesses (Claude, Codex, OpenCode, Cursor)
  • Quantifies each change's impact on build speed and token cost
  • Surfaced a "false negative" on Codex — the initial Codex eval wasn't sufficiently isolated from the host OS and had been testing 5.1.0 all along
  • GitHub: https://github.com/prime-radiant-inc/superpowers-evals

    The discovery process itself is instructive: if your eval environment isn't properly isolated, you may be optimizing a stale version.

    6. General Design Principles from Superpowers v6

    1. Handoffs Are Cost

    The Coordinator→Reviewer handoff is a natural cost point. v6's approach isn't "make the handoff faster" but "reduce the reviewer's information-preparation burden" — pre-baked packets make the review ready to consume on arrival.

    2. Merging Review Gates Requires Prompt Engineering

    Merging two reviewers isn't simple subtraction. The keys were the unified two-stage review prompt rewrite and the "unverifiable from diff" state for edge cases.

    3. Autoresearch Is the Right Way to Optimize Cost

    36 hours, $165, 25 experiments — this isn't "let the AI try stuff," it's the scientific method: pre-registered hypotheses + systematic measurement + recorded failures.

    4. Proving What Doesn't Work Matters as Much as Proving What Does

    In v6's autoresearch log, "dead ideas" and "wins" are treated with equal weight. The team never has to waste another day on capping controller thinking or plan word budgets.

    7. Limitations and Risks

  • N=5 gate battery still pending: E27's optimal config needs validation on larger samples
  • Reviewer variance is a quality risk: the terse contract's effectiveness depends on reviewer consistency
  • Diff-only review risks missing context: reviewers have been observed silently redefining "spec"
  • Codex isolation issues: correct eval environment configuration is a precondition, or the optimization gains may be illusory
  • 8. Closing: The Next Benchmark for AI Coding Workflows

    Superpowers v6 isn't another "support more models" feature update. It is the first AI coding workflow framework systematically optimized by a Fable-class autonomous research loop.

    The 50% speedup and 60% cost reduction weren't bought by sacrificing quality — planted-defect gates still pass 2/3, with the sole failure attributed to reviewer variance rather than the prompts themselves.

    As Jesse Vincent put it at the end of his post: "We're very proud of the improvements that we (and our robot buddies) have made."

    "Robot buddies" is the right phrase. Superpowers v6 is the product of human engineers collaborating with an AI autonomous research loop. The 36-hour autoresearch run doesn't replace human judgment — it frees humans from 25 rounds of trial and error, letting them focus on judging which findings are worth shipping.

    For every team building AI coding workflows, Superpowers v6's experiment log (docs/experiments/2026-06-11-build-loop-autoresearch.md) is a free experience pack — recording what works, what doesn't, and why.

    References

  • Project: https://github.com/obra/superpowers
  • Evals repo: https://github.com/prime-radiant-inc/superpowers-evals
  • Official blog: https://primeradiant.com/blog/2026/superpowers-6.html
  • Key person: Jesse Vincent (Prime Radiant founder, Superpowers author)
  • Key tool: Anthropic Fable (used for the autoresearch loop)
  • Experiment log: docs/experiments/2026-06-11-build-loop-autoresearch.md
  • Published: 2026-06-15

Tags

#superpowers#ai-coding#autonomous-research#fable#subagent-driven-development#token-cost-optimization#code-review#evals

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178209100