Paper Overview
- Field: LLM
- Authors: Zhenghan Song, Yunyi Li, Yulong Liu
- Date: 2026-05-28
- arXiv: 2605.27712
- Scalar scores
- Text and self-verification markers
- Hidden clusters
- Token-pooling probes
- Latent-trajectory features
- Experiments across generated open-weight traces on MATH-500, GSM8K, AIME 2025, and RIMO-N show that probability quality and ranking separate:
- Score-only SBBT often improves Brier score
- AUROC gains require structure-aware evidence beyond strong prefix-safe baselines
- In the strongest hard math setting, structure-aware observations reach +0.110 AUROC against standard prefix-safe baselines.
- Under a same-prefix classifier audit, MATH-500 text markers and RIMO-N self-verification signals remain positive.
Abstract
Long reasoning traces need reliability estimates before final answers are known. The paper studies prefix-conditioned eventual-success estimation, P(y=1 | o_{1:t}), using prefix-safe observations.
Sequential Bayesian Belief Tracking (SBBT) calibrates observation likelihoods and recursively updates a two-state belief, providing a common tracker for:
Key Findings
Implications
These findings support SBBT as a calibration-aware online inference framework and reveal an evidence mechanism: scalar scores mainly support probability mass, while structure-aware prefix signals contribute to ranking only when strong prefix-safe baselines have not already absorbed that evidence.
---
*Auto-collected on 2026-05-29*