The Overthinking Problem in Reasoning Models: How PUMA Learns to Stop at the Right Time
> Min, D. et al. *Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models.* arXiv:2605.17672, 2026. > University of Illinois Chicago, Google Research, Politecnico di Milano
---
1. An Embarrassing Finding: 41-52% of Reasoning Is Filler
If you observe the reasoning traces of DeepSeek-R1 or o1, you'll notice a pattern: the model finds the correct solution path at some step, but then spends dozens more steps saying things like "let me verify this result again…" or "to be safe, let me double-check…"
Researchers found that 41-52% of reasoning tokens are generated after the model has already found the correct answer. These tokens don't change the final answer or add new logical steps — they just re-chew conclusions already established. This is "overthinking": like a student who finishes the exam early but keeps re-checking correct answers until the bell rings.
---
2. Why Existing Early-Exit Methods Fall Short
Existing strategies fall into two camps:
- Answer confidence (e.g., DEER): stop when the model is confident in the answer.
- Answer consistency (e.g., Answer Convergence, Dynasor): stop when repeated probes yield the same answer.
- Prompt-based compression (CCoT, CoD, Plan-and-Budget): reduces tokens but accuracy collapses (e.g., CCoT drops Qwen3-30B from 81.7% to 54.3%) because hard budget limits cut necessary intermediate reasoning.
- Answer-level early exit (Answer Convergence, Dynasor, DEER): either too aggressive (accuracy crashes) or inconsistent in efficiency due to frequent answer probing.
- PUMA only considers stopping when reasoning is semantically redundant, so it probes answers far less often.
- Redundancy detector overhead: only 0.4-1.1% of runtime; answer probing adds 0.2-0.57 s per question.
- Wall-clock speedup: 1.40x on DS-7B, 1.28x on DS-14B on average. DEER and Dynasor are sometimes slower than full CoT.
- Answer confidence suffers from premature self-assurance: an early intermediate conclusion can be confident, wrong, or merely local.
- Answer consistency suffers from lucky guesses: the same answer appearing twice doesn't mean the derivation is complete.
- Semantic redundancy reflects the reasoning process itself: when consecutive steps restate established content ("let me verify… this is correct… no omissions…"), the reasoning has entered a converged state — regardless of how the answer looks.
- Min, D. et al. (2026). Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models. *arXiv:2605.17672*.
Both share a blind spot: **they monitor whether the *answer* is stable, not whether the *reasoning* has converged. A model may glimpse the answer at step 3 while the real derivation — auxiliary lines, equations, boundary checks — completes at step 10. Stopping at step 3 yields a correct-but-unsupported answer.
Reasoning typically moves through an exploration phase (low confidence/consistency), a self-correction phase (fluctuating signals), and a convergence phase (high confidence, stable answers, high semantic redundancy). Existing methods exit too early (confidence happens to spike during exploration) or too late (wasting repeated-verification tokens in convergence).
PUMA's core insight: monitor the semantic redundancy of reasoning steps themselves.
---
3. PUMA's Two-Layer Mechanism: "Should We Stop?" Then "Can We Stop?"
PUMA = Progress-aware Unified Monitoring framework for Adaptive early exit**. It decouples *considering* stopping from *confirming* stopping:
1. Redundancy Detector — a contrastively fine-tuned Qwen3-Embedding-0.6B computes embeddings per reasoning step and scores redundancy as max cosine similarity between the current step and the previous k steps (default k=1). A score above a threshold (0.35) means the step repeats earlier content — reasoning has converged. Contrastive training distinguishes genuinely progressive steps from re-verification and circular reasoning; generic semantic similarity is not enough.
2. Answer Verification — redundancy alone doesn't trigger exit. The framework induces a trial answer from the current prefix, computes its confidence (geometric mean token probability), and requires, across L consecutive redundancy points: the first probe's confidence above λ=0.98, all subsequent answers consistent, and no confidence drop beyond ε=0.03.
3. Loop Breaker — a fallback: after a minimum step count (50) with m consecutive redundant steps and a historical best-confidence answer above 0.8, exit with the best observed answer.
---
4. Results: 26.2% of Tokens Saved, Accuracy Intact
| Model | Baseline acc. | PUMA acc. | Token reduction | |-------|--------------|-----------|-----------------| | DeepSeek-R1-Distill-Qwen-7B | 60.8% | 60.2% | 35.6% | | DeepSeek-R1-Distill-Qwen-14B | 66.5% | 66.5% | 28.3% | | DeepSeek-R1-Distill-Qwen-32B | 71.8% | 71.8% | 22.1% | | Llama-3.1-Nemotron-Nano-8B | 54.3% | 54.3% | 29.4% | | Qwen3-30B-A3B-Thinking | 81.7% | 81.7% | 15.8% |
Average 26.2% token reduction with accuracy fully preserved. Occasionally accuracy even improves slightly — continued reasoning after finding the answer sometimes causes models to "second-guess" themselves into errors, which PUMA avoids.
Compared to Baselines
Reasoning Chain Quality
Judged by GPT-5.4-thinking, PUMA's preserved chains score best on coherence, conciseness, and argument sufficiency (completeness 95, coherence 90, conciseness 85, sufficiency 90) — it trims repetition rather than amputating the narrative.
Real Speedups
Zero-Shot Generalization
Without retraining or tuning: LiveCodeBench code generation saves 18-19% tokens (pass@1 change ≤1.5 points); MathVista/MathVision vision-language reasoning saves 23.8-33.6% tokens (accuracy change ≤1.5 points). Reasoning-level semantic redundancy is a domain- and modality-agnostic signal.
---
5. Key Insight: Why Semantic Redundancy Beats Answer Confidence
---
6. Limitations
1. The detector must be trained via contrastive learning, though it's lightweight (0.6B). Future work explores having models internalize stopping behavior during training. 2. Threshold tuning: the similarity threshold (0.35) was calibrated on held-out data; extreme cases may need adaptive adjustment.
3. Blind spot for non-redundant filler: PUMA detects repetition, not novel-but-irrelevant content (e.g., philosophical digressions or hallucinated tangents), which is semantically "new." 4. Multi-part convergence: with a local window (k=1), convergence of earlier sub-problems may go undetected, though redundancy typically appears in the final phase.
---
7. Conclusion
PUMA's essence isn't the early-exit mechanism itself, but a deeper realization: good reasoning isn't the longest — it's the one that stops when exploration is done and brakes when repetition begins. With 26.2% token savings, 1.4x speedup, and fully preserved accuracy, PUMA preserves the reasoning chain's value as an explanation: users see a complete, coherent, filler-free derivation rather than a truncated one.
---
Reference