You ask a language model to solve a math problem and it gets it wrong. You have two options:
A. Make it reflect — "You got that wrong. Think about why, and try again." B. Make it try again — same problem, new random seed, then take the majority answer.
Intuitively, A seems smarter. Reflection is central to human learning, and a model "examining its own output" sounds far more sophisticated. But a July 2026 paper, using 36 controlled experiments, shows that at equal token budgets, B almost always beats A—and by a significant margin.
Paper: *Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B* (arXiv:2607.28576)
Method Dividend vs. Budget Dividend
First, a confounding variable that's easy to overlook: almost every "reflection" method makes the model generate more tokens.
Self-Refine has the model answer, critique, and rewrite—three rounds of output. Reflexion has the model answer, reflect on errors, and re-answer with the reflection—also multi-round. Best-of-N generates N answers and picks the best—N times the output. Debate has multiple model instances argue against each other—even more exponential growth.
And generating more tokens by itself improves accuracy. Wang et al. 2024 already showed that sampling the same question multiple times and taking a majority beats single-shot sampling by a lot. This means: if a "reflection method" beats a single answer, you can't tell whether the idea of reflection is working, or whether it's just the extra tokens.
This paper did something simple that nobody had done before: equalize the token budgets. Let reflection methods and repeated sampling consume the same number of tokens, then compare accuracy.
36 Comparisons, Zero Wins
Experimental design:
- 7 methods: single-shot CoT, repeated sampling (majority vote), Self-Refine, Reflexion, Best-of-N (model picks), Best-of-N (count picks), debate
- 3 model scales: 1.5B, 3B, 7B
- 2 math benchmarks: GSM8K and MATH
- 150 problems, paired comparisons per problem
- Statistically rigorous: bootstrap confidence intervals + multiple-comparison corrections
- Count-based: generate 8 answers, pick the most frequent one
- Model-selected: generate 8 answers, let the model pick the best one
- Repeated sampling + majority voting is almost always the better choice
- The "intelligence" of model self-inspection may be a method dividend illusion—it looks sophisticated, but the actual gains come from the mundane fact of generating more tokens
- As models get bigger, the cost of "letting the model pick its own answer" decreases, but the benefit still doesn't exceed simple counting
The result in one sentence: no self-inspection method reliably beats repeated sampling at equal token budgets. Of 36 comparisons, 0 were significantly positive, 10 were significantly negative (all involving models inspecting their own outputs), and the remaining 26 were indistinguishable from zero.
Even more damning: all 18 "self-inspection" comparisons were negative. Self-Refine and forced Reflexion scored 3.6 to 10.1 percentage points below repeated sampling on the 7B model.
The Cost of "Choosing"
One detail is particularly interesting. Best-of-N has two selection modes:
Intuitively, letting the model choose should be better—it can understand the answers. But the experiments found: on the 1.5B model, count-based selection beat model selection by 8.0 and 11.3 percentage points. The model choosing its own answers was worse.
As models scale up (7B), the gap shrinks to 2.0 and 1.3 percentage points—no longer significant. Small models are bad at picking their own answers; larger models are slightly better, but still don't beat simple counting.
What does this tell us? Having a model examine its own output is inherently a costly operation. The model spends tokens judging, comparing, and selecting—tokens that would yield more if spent on additional samples.
Reflexion's Embarrassment
The most awkward case is Reflexion. Its design: the model answers, judges whether it was correct, and if wrong, reflects and retries.
But on the 1.5B model, Reflexion never triggered its retry mechanism—it judged itself correct every time, silently degenerating into single-shot CoT. This means the "Reflexion effect" reported in the original paper was, on small models, not reflection at work at all—just a single answer.
This is a classic "method dividend" illusion: you think the method is doing something, but it never even activated, and the improvement you saw came from somewhere else.
What This Means
The direction of "making models reflect" may need re-examination.
This doesn't mean reflection is useless—when token budgets aren't constrained, Self-Refine and Reflexion can improve accuracy. But if you care about per-token efficiency (which you should in production), then:
A Deeper Insight
This paper suggests a principle: when a method claims to "beat the baseline," first ask how many extra tokens it costs.
In the LLM era, tokens are compute cost. A method that spends 3x tokens for a 5% accuracy gain should not be published when a simple baseline spends 3x tokens for a 6% gain.
Academia has long ignored this "budget-equalized" control. One reason: "making models reflect" sounds sexy—it feeds a narrative of "AI becoming smarter"—while "sample a few more times and take a majority" sounds too plain.
The value of this paper isn't overturning a few methods; it's doing a neglected controlled experiment rigorously. The conclusion is merciless, but that's how science should be.
---
Paper link: https://arxiv.org/abs/2607.28576