English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Budgets

Forum topic · ✨步子哥 · 2026-08-01

Summary

A controlled study by Iliya Mirzaei (arXiv:2607.28576) compares self-reflection methods against naive repeated sampling with majority voting under strictly equal token budgets. Across 7 methods (Chain-of-Thought, Repeated Sampling, Self-Refine, Reflexion, Best-of-N with model selection or majority vote, Multi-Agent Debate), 3 model scales (1.5B, 3B, 7B), and 2 math benchmarks (GSM8K, MATH), none of 36 paired comparisons showed self-reflection stably beating the equal-cost repeated-sampling baseline, while 10 comparisons favored the baseline. Self-verification comparisons were negative in all 18 cases, and model-based selection underperformed majority voting by up to 11.3 points on smaller models. Notably, Reflexion never triggered its retry mechanism on the smallest model, silently reducing to single CoT. The findings suggest reported gains from self-reflection largely stem from extra generation rather than genuine reflection, and argue that Agentic AI papers must include cost-controlled baselines.

An Uncomfortable Conclusion for the Agent Community

Over the past two years, "letting the model reflect on itself" has been a default component of Agentic AI: Self-Refine, Reflexion, Multi-Agent Debate, Best-of-N with Self-Verification... The idea is always the same—have the model generate more text, then pick the best output.

This paper did something simple that no one had done seriously: compare again with token budgets strictly held equal.

The result is awkward: across 36 paired comparisons, no self-reflection method stably beat the naive baseline of "repeated sampling + majority voting" in any configuration. 10 comparisons showed self-reflection stably worse, and all 18 self-verification comparisons were negative.

Experimental Design

  • 7 methods: Chain-of-Thought (baseline), Repeated Sampling + Majority Vote, Self-Refine, Reflexion, Best-of-N with Model Selection, Best-of-N with Majority Vote, Multi-Agent Debate
  • 3 parameter scales: 1.5B, 3B, 7B
  • 2 math benchmarks: GSM8K, MATH
  • 150 questions per benchmark
  • Every token counted as cost—including critique, reflection, debate rounds, and verification tokens
  • Paired design: each method vs. cost-matched repeated sampling on the same questions, with bootstrap confidence intervals and multiple-comparison correction
  • Key Numbers

  • 36 paired comparisons: 0 showed self-reflection methods stably winning
  • 10 showed self-reflection stably worse—all were methods where the model checks its own output
  • 18/18 self-verification comparisons negative
  • Best-of-N "let the model pick" vs. "straight majority voting":
  • 1.5B: model selection trailed majority voting by 8.0 and 11.3 percentage points
  • 7B: lower by 2.0 and 1.3 points (no longer significant)
  • Self-Refine and forced Reflexion: still 3.6 to 10.1 percentage points below baseline at 7B
  • The darkest joke in the paper: Reflexion never triggered its own retry mechanism on the smallest model—it judged itself correct every time and silently degenerated into single-shot CoT. Quote from the paper: "It judged itself correct every time and silently became a single chain of thought."

What This Means

The gains attributed to "self-reflection" may never have come from reflection—they came from generating more times.

When a paper claims "Self-Refine improves over CoT by 12 points," the true source of the improvement is likely that Self-Refine generates 3-5x more tokens while single-shot CoT generates once. Equalize the token budget, and the magic of "reflection" disappears.

Harsher still: self-verification (having the model check its own output) is not just useless but harmful—the more the model checks, the worse it gets. This is another version of the same phenomenon as "Looping Is Not Reliability" (correctness is not an absorbing state; 16% of correct patches were reverted in the second round): self-verification doesn't just fail to fix errors—it flips correct answers into wrong ones.

The scale effect is also interesting: the smaller the model, the more harmful "let the model pick"; at larger scale the gap shrinks to insignificance. This suggests self-verification ability improves slowly with scale, but even at 7B it never turned into a positive gain.

A Methodological Warning

This paper sounds an alarm for the entire Agentic AI field: many papers claiming "agents beat baselines" may be comparing against strawmen—baselines whose costs were never controlled.

The correct practice: any paper claiming "method X works" must answer one question—"under an equal token budget, does X still beat repeated sampling?" If it cannot, X's effectiveness is unproven.

This connects perfectly with the "evaluation blind spot law": you optimize what you measure, and what isn't measured is where the problems hide. Since no one previously measured "cost-matched baselines," "self-reflection works" became the hiding place.

---

Paper title: Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Author: Iliya Mirzaei arXiv: https://arxiv.org/abs/2607.28576

Tags

#self-reflection#llm-agents#repeated-sampling#token-budget#evaluation#self-refine#reflexion#chain-of-thought

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503852