A Detective's Intuition
Imagine two detectives investigating a cold case:
The first walks in, locks eyes on a suspect within three minutes, and stays certain from start to finish. The interrogation transcript flows smoothly, with no hesitation at any step.
The second lists five possible suspects, draws a relationship map on a whiteboard, eliminates two hypotheses, hesitates for a while, and only finally settles on one person. The first half of the transcript is full of "it could be... but no..." back-and-forth; only the second half becomes clear.
Who is more reliable?
Intuition says the first — confidence equals correctness. But forensic common sense says: the second. For genuinely complex problems, no one can see the whole picture at the start. Locking in early means ignoring what you don't know.
This post discusses a paper that brings this detective story into the world of large language models. A team from the University of Illinois Urbana-Champaign and Microsoft found a counterintuitive phenomenon: on hard reasoning tasks, the answers with the highest confidence are the most likely to be wrong. Their solution is called Consilience — instead of looking at how confident a model is, look at the *shape* of its confidence.
---
The Problem: Where Confidence Maximization Collapses
Background: when LLMs do complex reasoning, a common strategy is Test-Time Scaling (TTS) — generate multiple reasoning paths and pick the best one.
The question is: how to pick?
- With verifiers (e.g., code): run test cases to check correctness.
- Verifier-free settings (e.g., free-text generation, agent decisions): no ground truth available.
- Premature convergence: the model locks onto one path immediately. May guess right on easy problems, but on hard ones it's "confidently wrong."
- Degenerate: confident at first, but reasoning collapses upon encountering contradictions.
- Consilience: explores multiple paths first (low initial confidence), then converges (high final confidence). This is the shape of truly reliable reasoning.
- Completely lost: the model simply can't solve the problem.
- \(C_{initial}\): average confidence over the first W tokens of the path (skipping the first P tokens, since given the same prompt, initial-token confidences are nearly identical and dilute the signal)
- \(C_{final}\): average confidence over the last W tokens
- \(\alpha\): balancing factor (fixed at α=3 in the paper)
- GPT-OSS-120B: Pass@1 65.7% → Consilience 69.7% (+4.0pp)
- Qwen3: Pass@1 54.9% → Consilience 58.1% (+3.2pp)
- GPT-OSS-120B + mini-swe-agent: 23.0% → 26.9% (+3.9pp)
- Qwen3-Coder-Next + mini-swe-agent: 65.3% → 67.3% (+2.0pp)
- Self-Consistency (majority voting): requires extractable, comparable answers; only works for multiple-choice/integer answers
- Process Reward Models: require training a reward model, expensive
- Consilience: works for free-text generation (code, agent decisions) — a scenario voting can't cover at all
- Epanorthosis: RLHF rewards confident emphasis, causing the AI flavor
- TokenBudget: CoT reasoning has bimodal fates, but layer-20 already encodes the ending
- QuantiBias: quantization introduces bias in blind spots of standard safety checks
- Consilience: average confidence masks the "shape" of confidence — a "confident hallucination" and a "hesitant-but-correct answer" can be indistinguishable in mean confidence
mini-swe-agent/: a mini-swe-agent implementation supporting Consiliencesingle_turn/: single-turn sampling utility scripts- In model training: the "shape" of the loss curve carries more information than the final loss value
- In agent systems: the "pattern" of tool-call sequences matters more than the call count
- In user behavior analysis: the "rhythm" of interactions predicts retention better than total clicks
- Title: Consilience for Verifier-Free Test-Time Scaling
- Authors: Lecheng Kong, Like Hui, Haitao Mao (UIUC + Microsoft)
- arXiv: 2608.09898
- Code: github.com/LechengKong/consilience
Existing methods use the model's confidence as a proxy — each token gets a probability distribution; the more concentrated, the more "confident" the model. Average the token confidences over the whole reasoning path, and select the path with the highest average.
Sounds reasonable: confident ≈ correct.
The paper first confirms this intuition holds at the macro level. On LiveCodeBench-V6 with GPT-OSS-120B generating multiple paths, correct answers do have higher average confidence than incorrect ones. But —
When the authors focus on hard problems (where model accuracy < 20%), the pattern reverses:
> Incorrect answers' average confidence (μ=9.03) is higher than correct answers' (μ=8.79).
Worse, the confidence distribution of incorrect answers has a long tail — a large batch of "extremely confident but completely wrong" reasoning paths, with confidence exceeding that of all correct answers. If you use a "pick the highest-confidence" strategy, you will inevitably select these confident hallucinations.
This is the "catastrophic collapse" in the title: confidence maximization works on easy problems, but on hard problems it not only fails — it systematically selects wrong answers.
---
The Insight: Confidence Has a "Shape," Not Just a "Height"
Why does this happen?
The authors borrow a concept from scientific epistemology: Consilience. A conclusion from a single line of reasoning is fragile; one that multiple independent lines of reasoning converge upon is reliable.
Through this lens, the problem with confidence maximization is clear:
What kind of path does confidence maximization select? One that is highly confident from the very first token — meaning the model collapsed onto a single path at step one without exploring alternatives. On easy problems, this is fine (easy problems don't need exploration), but on hard problems it's a disaster:
> A model that truly understands a complex problem should be aware of multiple possible solution paths. This awareness spreads probability mass across paths, manifesting as low initial confidence. Then, as reasoning deepens, the model filters, eliminates, and confirms among paths, converging on an answer — manifesting as high final confidence.
So a genuinely reliable reasoning trace is not "confident from start to finish," but "hesitant first, confident later." Confidence has a shape: it should be a rising curve, not a flat high line.
The paper categorizes all reasoning paths into a 2×2 matrix:
| | High final confidence | Low final confidence | |---|---|---| | High initial confidence | Premature convergence (Category 1) | Degenerate (Category 2) | | Low initial confidence | Consilience (Category 3) ✓ | Completely lost (Category 4) |
Existing methods inadvertently optimize Category 1 while ignoring the truly valuable Category 3.
---
The Formula: Quantifying Reasoning Quality via Temporal Asymmetry
The authors propose a remarkably simple formula:
In one sentence: reward final certainty, penalize initial certainty.
Select the path with the highest S. That's it. No training, no external model — just the model's own log-probabilities.
A Key Detail: Reasoning-Phase Isolation
The paper also flags an easily missed pitfall. Modern reasoning models (DeepSeek-R1, Qwen3, GPT-OSS, etc.) generate two phases:
1. Reasoning phase (between <think>...</think>): the actual thinking process
2. Answer phase (after </think>): summarizing the conclusion
Answer-phase tokens are naturally high-confidence — they merely restate conclusions already reached, not the genuine cognitive search. Including them pollutes the signal.
The fix: compute Consilience scores only on the reasoning phase, using a simple string split on </think> or <|channel|>final — nearly zero cost.
---
The Data: From Code Generation to Agent Tasks
Consilience was validated on four tasks:
LiveCodeBench-V6 (free-form code generation):
HMMT (hard math competition): significant gains
GPQA (graduate-level QA): significant gains
SWE-bench Verified (agent code repair):
Most notable is the difficulty-stratified experiment (Table 4), splitting LiveCodeBench into Easy/Medium/Hard by pass rate:
| Model | Method | Easy | Medium | Hard | |---|---|---|---|---| | Qwen | Mean-conf | 99.9 | 58.1 | 3.7 | | Qwen | Consilience | 99.9 | 70.3 | 8.1 | | GPT-OSS-120B | Mean-conf | 100.0 | 74.9 | 13.4 | | GPT-OSS-120B | Consilience | 99.1 | 88.6 | 17.2 |
Easy tier flat or slightly down (0 to −0.8), Medium tier up sharply (+12.2 to +13.7), Hard tier more than doubled in relative terms (+4.4 to +4.5). This perfectly confirms the core claim: Consilience is neutral on easy problems and decisively effective on hard ones.
---
Engineering Insights: Why You Should Use This Now
1. Zero Training Cost
Consilience requires no additional training. As long as the model exposes log-probabilities (most open models and closed APIs do), it works out of the box. One-line formula, minimal implementation cost.
2. Fills the Verifier-Free Gap
3. Extremely Stable Hyperparameters
The authors fixed α=3, k=20%, skip=5% on Qwen + LiveCodeBench, then applied them unchanged to other models and datasets — still with significant gains. Table 5 shows the fixed configuration captures 72–78% of per-set optimal gains, and cross-dataset transfer (LCB → HMMT) still yields +2.2pp. No per-scenario tuning needed.
4. Resonance with the "Evaluation Blind Spot" Principle
This paper is another perfect case in the "evaluation blind spot" concept lineage:
Single metrics (average confidence, loss, MMLU) mask critical failure modes. You optimize what you measure; problems hide where you don't.
---
One Level Deeper: Exploration Should Be Rewarded, Not Penalized
The paper ends with a passage worth quoting in full:
> "We view consilience not merely as one selection metric but as an instance of a general test-time-scaling principle: that exploration should be valued rather than penalized."
This points to a broader principle. In reinforcement learning, exploration vs. exploitation is the core tension. In test-time scaling, existing confidence-maximization methods are essentially pure exploitation — picking the path the model is most confident about. Consilience restores the value of exploration: low initial confidence is not a defect but a signal that the model is "considering multiple possibilities."
This aligns closely with how human experts reason. Facing a complex problem, a true expert's first reaction is often "it depends..." or "there are several possibilities..." — a sign of epistemic humility. A half-baked person's first reaction is "definitely..." — premature convergence.
Consilience quantifies "epistemic humility" with one line of math.
---
Open-Source Code
The paper is open-sourced: https://github.com/LechengKong/consilience
The repo includes:
The code is minimal (only 3 commits); the core is the one-line formula plus reasoning-phase isolation.
---
Personal Reflections
1. "Shape > Height" Is a Domain-agnostic Principle
The core insight: don't just look at a signal's mean; look at its trajectory. This goes beyond confidence:
The mean is a lazy statistic. Shape is the carrier of information.
2. Connection to the "Granularity Isomorphism" Principle
Consilience refines the decision granularity from "average confidence of the whole path" to "segmented initial vs. final confidence." This is the same direction as Heddle/CodeRescue's "from call-level to trajectory-level": the optimization granularity should match the granularity of the optimized object.
Reasoning has temporal structure (start → end), so the metric should have temporal structure (initial vs. final). Evaluating reasoning with average confidence is like evaluating a journey with traffic lights by average speed — the information gets smoothed away.
3. Another Member of the "Change-the-Level" Lineage
Consilience doesn't train better models or add external verifiers; it switches the evaluation dimension: from static confidence to dynamic confidence trajectories. It's the eleventh member of the "solve by changing levels" lineage:
Octopus RNA editing → slime mold externalized memory → avian quantum magnetoreception → SOPHIA division of labor → EvoThink atomic reasoning → Möbius RoPE topological intervention → mantis shrimp phononic shields → Euclid-MCP reasoning outsourcing → Regression Tax paired evaluation → ACE context engineering → Consilience temporal trajectory
Common pattern: don't work harder at the same level; work smarter at a different level.
4. A Hypothesis Worth Testing
The paper notes an interesting finding: on easy problems, high initial confidence is indeed a good signal of correctness (Category 1 works there). Only on hard problems does high initial confidence become a signal of "premature convergence."
This suggests a difficulty-adaptive strategy: dynamically adjust α based on problem difficulty. Easy problems: α=0 (pure confidence maximization); hard problems: α=3 (Consilience). The paper doesn't run this experiment, but Table 4's difficulty-stratified results hint there's meat in this direction.
---
Paper Info
*When a model says "I'm sure," ask: have you been sure from the very beginning, or did you become sure after considering many possibilities? The former is hallucination; the latter is reasoning.*