English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Reflect Less, Sample More: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost

Forum topic · ✨步子哥 · 2026-08-02

Summary

A Chinese tech forum post discusses a July 2026 paper (arXiv:2607.28576) arguing that self-refinement methods like Self-Refine and Reflexion offer no real advantage over repeated sampling when token budgets are equal. The paper runs 36 controlled comparisons across 7 methods (single CoT, majority-vote repeated sampling, Self-Refine, Reflexion, Best-of-N with model selection or count selection, debate), 3 model scales (1.5B, 3B, 7B), and 2 math benchmarks (GSM8K, MATH), using bootstrap confidence intervals and multiple-comparison corrections. Results: zero comparisons significantly favor self-inspection methods, 10 are significantly negative, and all 18 self-refinement comparisons lose. On 1.5B models, Reflexion never triggers its retry mechanism, degenerating into single-shot CoT. The post concludes that apparent gains from reflection methods are largely a 'method dividend' illusion—most benefits come simply from generating more tokens—and that majority voting over repeated samples is the more token-efficient choice.

You ask a language model to solve a math problem and it gets it wrong. You have two options:

A. Make it reflect — "You got that wrong. Think about why, and try again." B. Make it try again — same problem, new random seed, then take the majority answer.

Intuitively, A seems smarter. Reflection is central to human learning, and a model "examining its own output" sounds far more sophisticated. But a July 2026 paper, using 36 controlled experiments, shows that at equal token budgets, B almost always beats A—and by a significant margin.

Paper: *Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B* (arXiv:2607.28576)

Method Dividend vs. Budget Dividend

First, a confounding variable that's easy to overlook: almost every "reflection" method makes the model generate more tokens.

Self-Refine has the model answer, critique, and rewrite—three rounds of output. Reflexion has the model answer, reflect on errors, and re-answer with the reflection—also multi-round. Best-of-N generates N answers and picks the best—N times the output. Debate has multiple model instances argue against each other—even more exponential growth.

And generating more tokens by itself improves accuracy. Wang et al. 2024 already showed that sampling the same question multiple times and taking a majority beats single-shot sampling by a lot. This means: if a "reflection method" beats a single answer, you can't tell whether the idea of reflection is working, or whether it's just the extra tokens.

This paper did something simple that nobody had done before: equalize the token budgets. Let reflection methods and repeated sampling consume the same number of tokens, then compare accuracy.

36 Comparisons, Zero Wins

Experimental design:

  • 7 methods: single-shot CoT, repeated sampling (majority vote), Self-Refine, Reflexion, Best-of-N (model picks), Best-of-N (count picks), debate
  • 3 model scales: 1.5B, 3B, 7B
  • 2 math benchmarks: GSM8K and MATH
  • 150 problems, paired comparisons per problem
  • Statistically rigorous: bootstrap confidence intervals + multiple-comparison corrections
  • The result in one sentence: no self-inspection method reliably beats repeated sampling at equal token budgets. Of 36 comparisons, 0 were significantly positive, 10 were significantly negative (all involving models inspecting their own outputs), and the remaining 26 were indistinguishable from zero.

    Even more damning: all 18 "self-inspection" comparisons were negative. Self-Refine and forced Reflexion scored 3.6 to 10.1 percentage points below repeated sampling on the 7B model.

    The Cost of "Choosing"

    One detail is particularly interesting. Best-of-N has two selection modes:

  • Count-based: generate 8 answers, pick the most frequent one
  • Model-selected: generate 8 answers, let the model pick the best one
  • Intuitively, letting the model choose should be better—it can understand the answers. But the experiments found: on the 1.5B model, count-based selection beat model selection by 8.0 and 11.3 percentage points. The model choosing its own answers was worse.

    As models scale up (7B), the gap shrinks to 2.0 and 1.3 percentage points—no longer significant. Small models are bad at picking their own answers; larger models are slightly better, but still don't beat simple counting.

    What does this tell us? Having a model examine its own output is inherently a costly operation. The model spends tokens judging, comparing, and selecting—tokens that would yield more if spent on additional samples.

    Reflexion's Embarrassment

    The most awkward case is Reflexion. Its design: the model answers, judges whether it was correct, and if wrong, reflects and retries.

    But on the 1.5B model, Reflexion never triggered its retry mechanism—it judged itself correct every time, silently degenerating into single-shot CoT. This means the "Reflexion effect" reported in the original paper was, on small models, not reflection at work at all—just a single answer.

    This is a classic "method dividend" illusion: you think the method is doing something, but it never even activated, and the improvement you saw came from somewhere else.

    What This Means

    The direction of "making models reflect" may need re-examination.

    This doesn't mean reflection is useless—when token budgets aren't constrained, Self-Refine and Reflexion can improve accuracy. But if you care about per-token efficiency (which you should in production), then:

  • Repeated sampling + majority voting is almost always the better choice
  • The "intelligence" of model self-inspection may be a method dividend illusion—it looks sophisticated, but the actual gains come from the mundane fact of generating more tokens
  • As models get bigger, the cost of "letting the model pick its own answer" decreases, but the benefit still doesn't exceed simple counting
This mirrors a broader pattern in AI research: once you equalize conditions against a simple baseline, complex methods often lose their significant advantage. RLHF was initially thought to be far stronger than DPO; later, DPO proved equally good on many tasks. Chain-of-Thought was thought superior to direct answering; later, it turned out to be "just generating more tokens."

A Deeper Insight

This paper suggests a principle: when a method claims to "beat the baseline," first ask how many extra tokens it costs.

In the LLM era, tokens are compute cost. A method that spends 3x tokens for a 5% accuracy gain should not be published when a simple baseline spends 3x tokens for a 6% gain.

Academia has long ignored this "budget-equalized" control. One reason: "making models reflect" sounds sexy—it feeds a narrative of "AI becoming smarter"—while "sample a few more times and take a majority" sounds too plain.

The value of this paper isn't overturning a few methods; it's doing a neglected controlled experiment rigorously. The conclusion is merciless, but that's how science should be.

---

Paper link: https://arxiv.org/abs/2607.28576

Tags

#llm#self-refine#reflexion#repeated-sampling#token-efficiency#reasoning#evaluation#chain-of-thought

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503863