Paper Overview
- Field: NLP
- Authors: Lecheng Kong, Like Hui, Haitao Mao
- Published: 2026-08-12
- arXiv: 2508.05137
- Actively penalizes high initial confidence
- Strictly requires final certainty
Abstract
Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because we do not have access to such high-quality verifiers in many real-world applications.
Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and need minimal access to internal model states, making the methods highly flexible across models and tasks.
In this paper, the authors demonstrate a critical limitation of existing confidence-based VF-TTS methods: they can catastrophically collapse on complex tasks. An interesting observed phenomenon is that uniformly high confidence usually indicates failed exploration, tending toward confidently wrong answers.
Key Insight
Robust epistemic search requires a specific confidence trajectory pattern: methods should perform exploratory branching at the start (manifesting as low initial confidence) and converge to high final-confidence solutions.
Method: Consilience
To realize this insight, the authors introduce consilience, a novel selection framework that explicitly evaluates temporal asymmetry of confidence during reasoning. It is implemented via a combined metric that:
Results
Extensive experiments on graduate-level math problems and free-form code generation show that consilience effectively outperforms existing baselines, validating this novel perspective on completion confidence.
---
*Auto-collected on 2026-08-12.*