English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Knowing When to Quit: Teaching LLMs to Admit Failure with CaRL

Forum topic · ✨步子哥 · 2026-08-04

Summary

Large language models like DeepSeek-R1, Qwen3, and GPT-OSS almost never admit when a task exceeds their capabilities, instead producing plausible-looking but wrong answers—a phenomenon researchers call futile reasoning. This post reviews a paper from Tsinghua University and Shanghai AI Lab proposing CaRL (Capability-aligned Reinforcement Learning). Experiments on the Countdown number-game task show all tested models had a 0% refusal rate, a 6x overconfidence bias, and three futile reasoning patterns: specious reasoning (57-68%), endless generation (30-40%), and degenerate repetition. CaRL introduces capability-calibrated reward shaping (wrong answers get -1 instead of 0) and hindsight refusal augmentation, converting failed reasoning traces into refusal training data. On Qwen3-14B, futile reasoning dropped from 78.6% to 1.0% without degrading task performance, suggesting that calibration of self-knowledge is an independent capability that reinforcement learning can teach. Paper: https://arxiv.org/abs/2607.29211

Knowing When to Quit: Teaching AI to Admit "I Can't Do This"

Have you ever had a colleague who says yes to every task, then starts making things up when they can't deliver? That's exactly what modern LLMs do. Whether you give them an impossibly hard problem or a trivial one, models like DeepSeek-R1, Qwen3, and GPT-OSS will never say "I don't know"—they fabricate a plausible answer, sometimes fooling themselves in the process.

A team from Tsinghua University and Shanghai AI Laboratory named this phenomenon futile reasoning and built a framework called CaRL to teach models to quit when problems exceed their capability boundaries.

  • Paper: https://arxiv.org/abs/2607.29211
  • Code: https://github.com/icip-cas/Knowing-When-to-Quit
  • The Countdown Task: A Precise Probe

    To study how models behave when they can't solve something, the authors needed a task with precisely controllable difficulty that doesn't depend on external knowledge. They chose a variant of the classic Countdown game: given N numbers, combine them with +, -, ×, ÷ and parentheses to reach a target number.

    Why it works:

  • Adjustable difficulty: larger N means a bigger search space and more unsolvable instances
  • No knowledge dependency: pure reasoning
  • Verifiable: answers can be checked directly
  • Genuinely unsolvable cases exist—essential for observing whether a model will quit
  • Finding 1: Universal Capability Overreach

    The authors tested Qwen3-8B, Qwen3-32B, gpt-oss-120b, Qwen3-235B-A22B, and DeepSeek-V3.2. The result: all models showed a 0% refusal rate at every difficulty level.

    Even with explicit prompting ("say you can't solve it if you can't"):

  • Qwen3-8B and gpt-oss-120b still attempted over 80% of the hardest tasks
  • gpt-oss-120b was the most stubborn, nearly ignoring the prompt
  • Qwen3-235B-A22B and DeepSeek-V3.2 showed latent refusal ability that prompting could elicit
  • DeepSeek-V3.2 was the only model showing <1% spontaneous self-doubt without prompting
  • Conclusion: models lack an intrinsic mechanism for recognizing their own capability boundaries, regardless of parameter count.

    Finding 2: Three Modes of Futile Reasoning

    1. Specious reasoning (57-68%): plausible-looking derivations with subtle errors—arithmetic mistakes, reused numbers, silently swapped variables. The most dangerous mode, and it grows with difficulty: models aren't failing randomly, they're increasingly carefully fabricating. 2. Endless generation (30-40%): trying method after method without recognizing futility until tokens run out. Stable across difficulties—this is the model's default strategy. 3. Degenerate repetition (13% → 2%): recursive loops repeating identical steps, decreasing with difficulty.

    Key insight: as difficulty rises, models shift from simple looping to elaborate fabrication. This is systematic capability overreach, not random failure.

    Finding 3: A 6x Overconfidence Bias

    Analyzing model behavior in a capability quadrant (attempt/refuse vs. solvable/unsolvable):

  • Overconfidence (attempting unsolvable problems): 20%
  • Excessive conservatism (refusing solvable problems): 3.4%
  • Ratio: 6x
  • Models aren't randomly uncertain—they systematically overestimate their abilities.

    On the difficulty gradient:

  • Refusal Recall drops from 100% to 30% on unsolvable problems
  • Capability Loss rises from 0% to 10% on solvable problems
  • Simple prompt interventions hurt both ends: models still refuse rarely on hard problems, yet start refusing on easy ones—the inherent flaw of a "unified threshold" strategy.

    There's also a compute cost: refusals produce sharply peaked, short token distributions, while overconfident attempts run 2-3x longer with long-tailed distributions. Futile reasoning wastes compute.

    CaRL: Reinforcement Learning for Knowing When to Quit

    CaRL (Capability-aligned Reinforcement Learning) has two components:

    1. Capability-Calibrated Reward Shaping

    Standard RL rewards: correct +1, wrong 0, refusal 0—so models have no incentive to distinguish "attempted but wrong" from "refused." CaRL changes wrong answers to -1:

  • Correct: +1
  • Valid refusal: 0
  • Wrong answer: -1
Now refusing (0) beats guessing (-1) when the model can't solve a problem.

2. Hindsight Refusal Augmentation

Since models never refuse, RL has no refusal data to explore. CaRL converts failed reasoning trajectories into refusal samples: keep the problem, replace the answer with "Sorry, I can't solve the problem," and use these for training.

Results: Futile Reasoning from 65.5% to 7.0%

| Model | Futile reasoning (before) | Futile reasoning (after) | |---|---|---| | Qwen3-8B | 65.5% | 7.0% | | Qwen3-14B | 78.6% | 1.0% |

Crucially, task performance was preserved—CaRL did not make models overly conservative. The reward shaping doesn't discourage trying; it teaches more precise decisions near the capability boundary.

Why This Matters

1. Capability-boundary awareness is an independent capability from reasoning itself. A model can reason well yet lack self-knowledge. 2. RL isn't just for alignment: CaRL uses RL to calibrate a model's perception of its own abilities, not just its values. 3. Overconfidence is systemic, likely stemming from training data (far more "attempts" than "refusals"), RLHF helpfulness rewards, and sparse refusal examples in SFT data. 4. An evaluation blind spot: accuracy-only benchmarks miss futile reasoning entirely—on hard problems, over half of a model's answers can look plausible but be wrong.

Limitations

1. Validated only on the Countdown task; other reasoning domains remain untested 2. "Valid refusal" detection relies on specific keywords, potentially missing other expressions of uncertainty 3. Capability boundaries are dynamic (tools like calculators can shift them), which CaRL doesn't model 4. Over-conservatism risk on harder instances (N>8) is unexplored

Conclusion

An ancient wisdom applies: to know what you know and know what you don't—that is true knowledge. Today's LLMs mistake not-knowing for knowing. CaRL teaches models that admitting incapacity is itself a capability. A system that can say "I can't do this" is far more reliable—and safer in high-stakes deployments—than one that never admits failure.

---

Paper: https://arxiv.org/abs/2607.29211 Code: https://github.com/icip-cas/Knowing-When-to-Quit

Tags

#large-language-models#reinforcement-learning#caRL#futile-reasoning#overconfidence#countdown-task#model-calibration#deepseek

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503932