Overview
This forum post discusses "Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs" by Izhar Ali (Rowan University), presented at EIML@ICML 2026 (2nd Workshop on Epistemic Intelligence in Machine Learning).
- Paper: arXiv:2607.20464
- Venue: EIML@ICML 2026
- Per-question uncertainty: disagreement across repeated samples on one question — like flipping a coin repeatedly to estimate its bias.
- Cross-question structure: correlated performance patterns across related questions, revealing capability dimensions (e.g., "weak at all problems requiring spatial reasoning").
- Group A: at most one eigenvalue beyond the noise floor, with a negligible TW z-score (≤ +0.05). Repeated sampling contains almost no cross-question structure.
- Group B: four clear signal eigenvalues, validated against a difficulty-matched Bernoulli null (500 Monte Carlo draws) which produces at most one.
The Central Question
The widely used self-consistency technique (Wang et al., 2023) generates multiple answers from one LLM and takes a majority vote, forming the basis of answer-repetition confidence, semantic entropy, and hallucination detection. All these methods assume that sampling variation carries information about what the model knows.
The paper asks: is that variation structured *knowledge signal*, or just *noise*?
Two Kinds of Uncertainty
Method: Marchenko-Pastur Random Matrix Theory
Ali builds binary correctness matrices (attempts × 500 questions) and applies Marchenko-Pastur (MP) theory: a purely random matrix has a characteristic eigenvalue spectrum, so any eigenvalue beyond the MP edge signals non-random structure. The Tracy-Widom (TW) statistic quantifies significance.
Experimental Design and Results
Group A (within-model): Qwen2.5-7B, 100 runs at temperature=1 → 100×500 matrix. Group B (cross-model): 24 different LLMs, one run each at temperature=0 → 24×500 matrix.
Why Temperature Sampling Is "Shallow"
Temperature uniformly scales logits: it does not differentiate domains, question types, or "I know this" vs. "I'm guessing." Variation from temperature is therefore unstructured noise. The post's analogy: sampling wanders different corridors of the *same* library; different models are entirely different libraries.
A striking finding: when Qwen2.5-7B answered a borderline question wrong, it chose the same wrong answer 87.4% of the time — "locally committed, globally incoherent." Errors are isolated events, not structured symptom clusters as in human diagnostic error patterns.
Practical Implications
1. Selective prediction: self-consistency works well for per-question confidence, but cannot identify which *categories* of problems need retraining. 2. Model ensembles: a 24-model ensemble is not a pricier version of 100 samples — it is a fundamentally different information source. 3. Hallucination detection: semantic-entropy-style methods remain valid for per-question uncertainty but miss deeper cross-question systematic hallucination patterns.
Philosophical Takeaway
Consistency across samples may be mere "local inertia" rather than understanding. Genuine knowing would show predictable, structured performance across related problems. The dimensionality gap suggests current LLM sampling behavior may resemble memory rather than understanding — isolated facts without deep structural association.
As the post concludes: self-consistency is not a cheap substitute for an ensemble — it is a fundamentally different, informationally poorer approximation. Like one instrument recorded 100 times versus a full orchestra, true intelligence may lie in how multiple voices resonate.
References
1. Ali, I. (2026). *Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs*. arXiv:2607.20464. EIML@ICML 2026. 2. Wang, X., et al. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. *ICLR*. 3. Lakshminarayanan, B., et al. (2017). Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. *NeurIPS*. 4. Marčenko, V.A., & Pastur, L.A. (1967). Distribution of eigenvalues for some sets of random matrices. *Mathematics of the USSR-Sbornik*. 5. Farquhar, S., et al. (2024). Detecting Hallucinations in Large Language Models Using Semantic Entropy. *Nature*.