> Note: This is a GEO-optimized version of the original topic, with a question-driven title, structured data, and FAQ to improve AI engine citation.
One-sentence takeaway: This post analyzes the core finding and engineering implications of "reflect less, sample more: the method advantages of Self-Refine and Reflexion may not exist."
Suppose you ask a language model a math question and it gets it wrong. You have two options:
A. Make it reflect — "You got that wrong. Think about why and try again." B. Make it try again — Same question, new random seed, then take the majority answer.
Intuitively, A seems smarter. Reflection is central to human learning, and a model "examining its own output" sounds far more sophisticated. But a new paper from July 2026, based on 36 controlled comparisons, shows: at equal token budgets, B almost always beats A, often by a wide margin.
Paper: *Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B* (arXiv:2607.28576)
Method Advantage vs. Budget Advantage
First, a confounding variable that is easy to overlook: almost all "make the model reflect" methods cause the model to generate more tokens.
Self-Refine has the model answer, then critique, then rewrite—three rounds of output. Reflexion has the model answer, reflect on errors, then re-answer with the reflection—also multi-round. Best-of-N generates N answers and picks the best—N times the output. Debate has multiple model instances argue against each other—exponential growth.
And generating more tokens by itself improves accuracy. Wang et al. 2024 already showed that sampling the same question multiple times and taking a majority beats a single sample by a lot. This means—if a "reflection method" beats a single answer, you cannot tell whether the "reflection" idea is working or whether it's just "generating more tokens."
This paper does something simple that nobody had done: equalize the token budget. Let reflection methods and repeated sampling consume the same number of tokens, then compare accuracy.
36 Comparisons, Zero Wins
Experimental design:
- 7 methods: single-pass CoT, repeated sampling (majority vote), Self-Refine, Reflexion, Best-of-N (model picks), Best-of-N (majority count), debate
- 3 model scales: 1.5B, 3B, 7B
- 2 math benchmarks: GSM8K and MATH
- 150 questions, pairwise comparisons per question
- Statistical rigor: bootstrap confidence intervals + multiple-comparison correction
- Count-based: generate 8 answers, pick the most frequent one
- Model-picked: generate 8 answers, let the model choose the best
- Repeated sampling + majority voting is almost always the better choice
- The "intelligence" of model self-review may be a method-advantage illusion—it looks sophisticated, but the actual gains come from the mundane fact of generating more tokens
- As models get larger, the cost of "letting the model pick its own answer" decreases, but it still doesn't beat simple counting
- Method advantage vs. budget advantage
- 36 comparisons, zero wins for reflection
- The cost of "letting the model choose"
The result in one sentence: no self-review method reliably beat repeated sampling at equal token budgets. Of the 36 comparisons, 0 were significantly positive, 10 were significantly negative (all methods that make the model review its own output), and the remaining 26 were indistinguishable from zero.
More damning: all 18 "self-review" comparisons were negative. Self-Refine and forced Reflexion trailed repeated sampling by 3.6 to 10.1 percentage points on the 7B model.
The Cost of "Choosing"
One detail is especially interesting. Best-of-N has two selection modes:
Intuitively, model-picked should be better—the model can understand the answers. But the experiments found: on the 1.5B model, count-based selection beat model-picked by 8.0 and 11.3 percentage points. Letting the model choose was worse.
As the model grows (7B), the gap shrinks to 2.0 and 1.3 percentage points—no longer significant. Small models are bad at picking their own answers; larger models are slightly better but still don't beat simple counting.
What does this mean? Having a model examine its own output is inherently a costly operation. The model spends tokens judging, comparing, and selecting—tokens that, if spent on more samples, yield greater returns.
Reflexion's Awkwardness
The most awkward result concerns Reflexion. The method is designed so that after answering, the model judges whether it was correct and, if wrong, reflects and retries.
But on the 1.5B model, Reflexion never triggered its own retry mechanism—it judged itself correct every time, silently degrading into single-pass CoT. This means the "Reflexion effect" reported in the original paper was, on small models, not reflection at work at all—just a single answer.
This is a classic "method advantage" illusion: you think the method is doing something, but the method never actually engaged—the gains you observed came from elsewhere.
What This Means
The "make the model reflect" direction may need to be re-examined.
This doesn't mean reflection is useless—when token budgets are unconstrained, Self-Refine and Reflexion do improve accuracy. But if you care about per-token efficiency (which you should in production), then:
This parallels a more general phenomenon in AI research: once complex methods are compared against simple baselines under equalized conditions, they often show no significant advantage. RLHF was initially thought to be far stronger than DPO, until DPO proved equally good on many tasks. Chain-of-Thought was thought to beat direct answering, until it turned out the gains came from "generating more tokens."
A Deeper Insight
This paper suggests a principle: when a method claims to "beat a baseline," first ask how many extra tokens it spends.
In the LLM era, tokens are compute cost. A method that spends 3x tokens for a 5% accuracy gain should not be published over a simple baseline that spends 3x tokens for a 6% gain.
But academia has long ignored this "budget-equalized" control. One reason: "making the model reflect" sounds sexy—it implies a narrative of "AI becoming smarter"—while "sample a few more times and take a majority" sounds too plain.
This paper's value is not overturning a few methods, but doing the neglected control experiment rigorously. The conclusion is merciless, but that's how science should be.
---
Paper link: https://arxiv.org/abs/2607.28576
FAQ
Q1: Who is this content for?
Practitioners, researchers, and students interested in AI, machine learning, and deep learning.
Q2: What are the key points?
See the links in the article.