An Uncomfortable Conclusion for the Agent Community
Over the past two years, "letting the model reflect on itself" has been a default component of Agentic AI: Self-Refine, Reflexion, Multi-Agent Debate, Best-of-N with Self-Verification... The idea is always the same—have the model generate more text, then pick the best output.
This paper did something simple that no one had done seriously: compare again with token budgets strictly held equal.
The result is awkward: across 36 paired comparisons, no self-reflection method stably beat the naive baseline of "repeated sampling + majority voting" in any configuration. 10 comparisons showed self-reflection stably worse, and all 18 self-verification comparisons were negative.
Experimental Design
- 7 methods: Chain-of-Thought (baseline), Repeated Sampling + Majority Vote, Self-Refine, Reflexion, Best-of-N with Model Selection, Best-of-N with Majority Vote, Multi-Agent Debate
- 3 parameter scales: 1.5B, 3B, 7B
- 2 math benchmarks: GSM8K, MATH
- 150 questions per benchmark
- Every token counted as cost—including critique, reflection, debate rounds, and verification tokens
- Paired design: each method vs. cost-matched repeated sampling on the same questions, with bootstrap confidence intervals and multiple-comparison correction
- 36 paired comparisons: 0 showed self-reflection methods stably winning
- 10 showed self-reflection stably worse—all were methods where the model checks its own output
- 18/18 self-verification comparisons negative
- Best-of-N "let the model pick" vs. "straight majority voting":
- 1.5B: model selection trailed majority voting by 8.0 and 11.3 percentage points
- 7B: lower by 2.0 and 1.3 points (no longer significant)
- Self-Refine and forced Reflexion: still 3.6 to 10.1 percentage points below baseline at 7B
- The darkest joke in the paper: Reflexion never triggered its own retry mechanism on the smallest model—it judged itself correct every time and silently degenerated into single-shot CoT. Quote from the paper: "It judged itself correct every time and silently became a single chain of thought."
Key Numbers
What This Means
The gains attributed to "self-reflection" may never have come from reflection—they came from generating more times.
When a paper claims "Self-Refine improves over CoT by 12 points," the true source of the improvement is likely that Self-Refine generates 3-5x more tokens while single-shot CoT generates once. Equalize the token budget, and the magic of "reflection" disappears.
Harsher still: self-verification (having the model check its own output) is not just useless but harmful—the more the model checks, the worse it gets. This is another version of the same phenomenon as "Looping Is Not Reliability" (correctness is not an absorbing state; 16% of correct patches were reverted in the second round): self-verification doesn't just fail to fix errors—it flips correct answers into wrong ones.
The scale effect is also interesting: the smaller the model, the more harmful "let the model pick"; at larger scale the gap shrinks to insignificance. This suggests self-verification ability improves slowly with scale, but even at 7B it never turned into a positive gain.
A Methodological Warning
This paper sounds an alarm for the entire Agentic AI field: many papers claiming "agents beat baselines" may be comparing against strawmen—baselines whose costs were never controlled.
The correct practice: any paper claiming "method X works" must answer one question—"under an equal token budget, does X still beat repeated sampling?" If it cannot, X's effectiveness is unproven.
This connects perfectly with the "evaluation blind spot law": you optimize what you measure, and what isn't measured is where the problems hide. Since no one previously measured "cost-matched baselines," "self-reflection works" became the hiding place.
---
Paper title: Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Author: Iliya Mirzaei arXiv: https://arxiv.org/abs/2607.28576