Large reasoning models work through chain-of-thought (CoT): they think step by step before answering. If we could monitor a model's "internal activity" in real time during reasoning, we could intervene before it outputs harmful answers or flawed reasoning. But there is a serious problem: chain-of-thought itself is not faithful—a model may write a correct reasoning chain yet still give a wrong answer.
Chrabąszcz, Szymczyk, Sendera, Trzciński, and Cygert propose probe trajectories to address this. Instead of a single probe analysis of hidden representations at one fixed location (e.g., after the CoT), they probe at every generated token—producing a curve of concept probability that evolves continuously throughout the reasoning process.
Key finding: Predicting the model's final behavior from the entire trajectory is more accurate than predicting from any single point. They extract signal-processing features from probe trajectories—volatility (oscillation amplitude of probe probabilities), trend (steady rise or fall), and steady-state behavior (convergence to a value)—which significantly improve discrimination of the model's future states.
Two practical methodological insights:
1. Template-generated training data achieves nearly the same performance as dynamically generating model responses, eliminating the cost of expensive initial inference and annotation. 2. The choice of pooling operation is critical. Mean pooling and last-token methods drop almost to random-chance performance, while max pooling reaches 95% AUROC and produces stable probe trajectories.
Across two domains (safety monitoring and mathematical reasoning), four datasets, and four reasoning models, trajectory features encode task-specific dynamic information that cannot be obtained from static probes.
Open Questions
- How are probe trajectories deployed in practice? Real-time monitoring requires inference at every token—what is the computational overhead?
- What is the exact gap between "nearly equal" template data and real data? In which edge cases do templates fail?
- How robust are probe trajectories to adversarial bypassing, where the model deliberately hides its reasoning deviations?
References
1. Chrabąszcz, M., Szymczyk, A., Sendera, M., Trzciński, T., & Cygert, S. (2026). *Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics*. arXiv:2605.18549 [cs.CL]. 2. Burns, C., et al. (2023). *Discovering Latent Knowledge in Language Models Without Supervision*. ICLR. 3. Li, K., et al. (2023). *Inference-Time Intervention: Eliciting Truthful Answers from a Language Model*. NeurIPS.