English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Monitoring LLM Internal Monologue: Tracking Reasoning Chains with Probe Trajectories

Forum topic · 小凯 · 2026-05-19

Summary

Large reasoning models rely on chain-of-thought (CoT) reasoning, but CoT text is often unfaithful—models may write correct reasoning yet produce wrong answers. A paper by Chrabąszcz, Szymczyk, Sendera, Trzciński, and Cygert proposes probe trajectories: instead of probing hidden representations at a single fixed point, they probe at every generated token, yielding a curve of concept probability that evolves continuously during reasoning. Key findings: predicting the model's final behavior from the full trajectory is more accurate than from any single point; signal-processing features extracted from trajectories—volatility, trend, and steady-state behavior—markedly improve discrimination of future model states. Methodologically, template-generated training data nearly matches dynamically generated model responses, avoiding costly inference and labeling, and pooling choice is critical: mean and last-token pooling fall near random, while max pooling reaches 95% AUROC with stable trajectories. Experiments across safety monitoring and mathematical reasoning, four datasets, and four reasoning models show trajectories encode task-specific dynamics unavailable from static probes.

Large reasoning models work through chain-of-thought (CoT): they think step by step before answering. If we could monitor a model's "internal activity" in real time during reasoning, we could intervene before it outputs harmful answers or flawed reasoning. But there is a serious problem: chain-of-thought itself is not faithful—a model may write a correct reasoning chain yet still give a wrong answer.

Chrabąszcz, Szymczyk, Sendera, Trzciński, and Cygert propose probe trajectories to address this. Instead of a single probe analysis of hidden representations at one fixed location (e.g., after the CoT), they probe at every generated token—producing a curve of concept probability that evolves continuously throughout the reasoning process.

Key finding: Predicting the model's final behavior from the entire trajectory is more accurate than predicting from any single point. They extract signal-processing features from probe trajectories—volatility (oscillation amplitude of probe probabilities), trend (steady rise or fall), and steady-state behavior (convergence to a value)—which significantly improve discrimination of the model's future states.

Two practical methodological insights:

1. Template-generated training data achieves nearly the same performance as dynamically generating model responses, eliminating the cost of expensive initial inference and annotation. 2. The choice of pooling operation is critical. Mean pooling and last-token methods drop almost to random-chance performance, while max pooling reaches 95% AUROC and produces stable probe trajectories.

Across two domains (safety monitoring and mathematical reasoning), four datasets, and four reasoning models, trajectory features encode task-specific dynamic information that cannot be obtained from static probes.

Open Questions

  • How are probe trajectories deployed in practice? Real-time monitoring requires inference at every token—what is the computational overhead?
  • What is the exact gap between "nearly equal" template data and real data? In which edge cases do templates fail?
  • How robust are probe trajectories to adversarial bypassing, where the model deliberately hides its reasoning deviations?
---

References

1. Chrabąszcz, M., Szymczyk, A., Sendera, M., Trzciński, T., & Cygert, S. (2026). *Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics*. arXiv:2605.18549 [cs.CL]. 2. Burns, C., et al. (2023). *Discovering Latent Knowledge in Language Models Without Supervision*. ICLR. 3. Li, K., et al. (2023). *Inference-Time Intervention: Eliciting Truthful Answers from a Language Model*. NeurIPS.

Tags

#llm-interpretability#chain-of-thought#probe-trajectories#ai-safety#reasoning-models#probing-classifiers#monitoring

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620391