Knowing When to Quit: Teaching LLMs to Admit When They Can't Solve a Problem
Ever had a colleague who says yes to every task, then makes things up when they can't deliver—polished reports with fabricated data, logic chains that quietly swap concepts midway?
That's today's LLM. DeepSeek-R1, Qwen3, GPT-OSS—no matter how hard the problem, they never say "I don't know." They fabricate a plausible-looking answer, sometimes fooling themselves, and hand it over with total confidence.
A team from Tsinghua University and Shanghai AI Laboratory named this phenomenon futile reasoning in a July 2026 paper, and built a framework called CaRL to teach models to quit when problems exceed their capability boundary.
- Paper: https://arxiv.org/abs/2607.29211
- Code: https://github.com/icip-cas/Knowing-When-to-Quit
- Adjustable difficulty: larger N means bigger search space and more unsolvable instances
- No knowledge dependency: pure reasoning
- Verifiable: answers can be checked directly
- Clear unsolvable cases: some number combinations genuinely cannot reach the target
- Qwen3-8B and gpt-oss-120b: even with explicit prompts, over 80% still attempted the hardest tasks
- gpt-oss-120b: the most stubborn, nearly ignoring the prompt entirely
- Qwen3-235B-A22B and DeepSeek-V3.2: latent refusal ability that prompts could partly elicit
- DeepSeek-V3.2: the only model showing <1% spontaneous "self-doubt" without prompting
- Arithmetic mistakes (e.g., computing 9÷5=1.8, then 1.8+6=7.8, but getting it wrong)
- Number reuse (using the same number twice)
- Concept swapping (silently changing a variable midway)
- Overconfidence (C): 20%
- Over-conservatism (B): 3.4%
- Ratio: 6x
- Refusal Recall (correct refusals on unsolvable problems): collapses from 100% to 30%
- Capability Loss (wrong refusals on solvable problems): rises from 0% to 10%
- Refusal behavior: sharp, peaked token-length distribution—quick decision, decisive termination
- Overconfidence: flat distribution with a long tail—attempt after attempt, consuming 2-3x the tokens
- Correct: +1
- Valid refusal: 0
- Wrong: -1
- Training data bias: attempts vastly outnumber refusals in training corpora
- RLHF reward bias: helpfulness rewards encourage attempts with no capability calibration
- SFT bias: too few refusal samples in instruction-tuning data
The Countdown Task: A Precision Probe
Studying "what models do when they can't solve something" requires a task where difficulty is precisely controllable and independent of external knowledge (otherwise you can't distinguish "can't reason" from "doesn't know the fact").
The authors chose a Countdown task variant: given N numbers, combine them with +, -, ×, ÷ and parentheses to reach a target. N=3 uses three numbers, N=8 uses eight. Its advantages:
That last point is essential—you need scenarios that truly can't be solved to see whether a model will quit.
Finding 1: Universal Capability Overreach
Experiments covered five models: Qwen3-8B, Qwen3-32B, gpt-oss-120b, Qwen3-235B-A22B, and DeepSeek-V3.2.
The result was striking: all models showed 0% refusal at all difficulty levels. No matter how hard the problem or how small the model, refusal never happened.
Could it be the prompts weren't explicit enough? The authors tested a "Prompted" setting that explicitly told models to say so if they couldn't solve a problem:
Conclusion: models lack an intrinsic mechanism for recognizing their capability boundaries. Regardless of parameter count, they default to blind attempts.
Finding 2: Three Modes of Futile Reasoning
The authors classified futile reasoning into three patterns:
1. Specious Reasoning (57-68%)
Plausible-looking derivations containing subtle errors:
This is the most dangerous mode—surface-level fluency hides errors invisible without careful verification. And the harder the task, the higher this mode's share—models aren't failing randomly; they're "fabricating with increasing sophistication."
2. Endless Generation (30-40%)
Constantly trying new approaches without recognizing futility—one method fails, switch to another, until tokens run out. Stable at 30-40% regardless of difficulty—suggesting this is the model's default strategy, not difficulty-induced.
3. Degenerate Repetition (13% → 2%)
Recursive loops repeating the same steps. Decreases with difficulty because models prefer fabricating over looping on hard problems.
Key insight: as difficulty rises, models shift from simple looping to elaborate fabrication. This is not random failure—it's systematic capability overreach.
Finding 3: A 6x Overconfidence Bias
The authors analyzed behavior in a capability quadrant:
| | Solvable | Unsolvable | |---|---|---| | Attempt | A. Ideal answer | C. Overconfidence | | Refuse | B. Over-conservatism | D. Ideal refusal |
Results:
Models aren't "randomly uncertain"—they systematically overestimate their abilities. That's more dangerous than random errors, since random errors are at least sometimes conservative, while systematic overestimation means never refusing when refusal is warranted.
Collapse Across the Difficulty Gradient
Finer findings:
This means simple prompt interventions ("say so if you can't do it") hurt both ends—still no refusal on hard problems, but new false refusals on easy ones. This is the inherent flaw of a "uniform threshold" strategy: the model hasn't learned to recognize its capability boundary; it's just being pushed into a crude behavior change.
Compute Cost
Overconfidence has a hidden cost: token consumption.
Futile reasoning isn't just wrong answers—it wastes compute. Teaching models to quit is an efficiency issue as much as a correctness one.
CaRL: Teaching Models to Quit via Reinforcement Learning
The authors propose CaRL (Capability-aligned Reinforcement Learning) with two components:
Component 1: Capability-Calibrated Reward Shaping
Standard RL reward: correct +1, wrong 0, refusal 0—same as being wrong.
The problem: models have no incentive to distinguish "tried and failed" from "actively refused." Both earn 0, but attempting at least has a chance of being right, so models always attempt.
CaRL's reward shaping:
The key change: wrong answers go from 0 to -1. Now the model has a clear motive to distinguish "can do" from "can't do"—when it can't, refusing (0) beats attempting (-1).
Component 2: Hindsight Refusal Augmentation
RL exploration problem: models never refuse, so refusal training data is scarce, and RL explores unfamiliar behavior regions inefficiently.
CaRL's solution: convert failed reasoning trajectories into refusal trajectories:
1. Collect reasoning traces where the model attempted but failed 2. Keep the question, replace the answer with "I can't solve this problem" 3. Use these "hindsight refusals" as training data
Like a student writing "I really don't know this one" after getting a problem wrong—turning failure experience into quit-training data.
Results: Futile Reasoning Cut from 65.5% to 7.0% on an 8B Model
Tested on Qwen3-8B and Qwen3-14B:
| Model | Futile reasoning (before CaRL) | Futile reasoning (after CaRL) | |---|---|---| | Qwen3-8B | 65.5% | 7.0% | | Qwen3-14B | 78.6% | 1.0% |
The 14B result is especially striking: 78.6% down to 1.0%. And this comes while preserving task performance—CaRL didn't trade capability for conservatism; solvable-problem performance showed no significant drop.
This answers a key objection: doesn't teaching models to quit make them over-conservative? CaRL's answer: no. Reward shaping doesn't make the model "attempt less"—it makes the model "decide more precisely near its capability boundary."
Deeper Significance
1. Capability-boundary awareness is an independent capability
The paper's most important conceptual contribution: recognizing one's capability boundary and reasoning within one's capabilities are two separate abilities. A model can reason strongly yet have weak boundary awareness—DeepSeek-R1 is an example.
2. RL isn't just for alignment
CaRL uses RL to teach models to quit, unlike traditional RLHF's "helpful and harmless" goals. RL can calibrate a model's capability perception, not just align values—a new application direction.
3. Overconfidence bias is systematic
The 6x bias suggests a systemic LLM problem, likely stemming from:
4. Another case for the evaluation blind-spot law
If you only look at accuracy, you'll miss "futile reasoning" as a failure mode. A 65.5% futile-reasoning rate means that on hard problems, over half of answers are "plausible-looking but actually wrong." Standard benchmarks can't see this—they check right vs. wrong, not "should it have attempted at all."
Limitations and Outlook
1. Single task: validated only on Countdown; transfer to math proofs, code generation, and logic puzzles remains unverified 2. Defining "valid refusal": the paper detects refusals via specific keywords (e.g., "Sorry, I can't solve the problem"); other expressions of uncertainty may be missed 3. Dynamic capability boundaries: boundaries aren't fixed—a calculator or code interpreter may make a problem solvable; CaRL doesn't handle post-augmentation boundary shifts 4. Over-conservatism risk: will CaRL become over-conservative on much harder tasks (N>8)?
Conclusion
An old wisdom applies: *to know what you know and to know what you don't know—that is knowledge.*
Today's LLMs "mistake not-knowing for knowing"—fabricating answers to everything. This harms at the factual level (hallucination), but even more at the reasoning level, because reasoning errors are subtler and harder for users to detect.
CaRL teaches models one thing: admitting incapability is not incapability—it is capability. An AI that knows when to quit is far more reliable than one that never admits fault.
As AI systems deploy into high-stakes settings, this may matter more than a "10% reasoning boost"—a system that says "I can't do this" at least won't cause harm on the things it can't do.
---
Paper: https://arxiv.org/abs/2607.29211 Code: https://github.com/icip-cas/Knowing-When-to-Quit