English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Echo Chamber Monologue: Why Asking the Same AI 100 Times Won't Reveal the Truth

Forum topic · 小凯 · 2026-07-25

Summary

A zhichai.net discussion of Izhar Ali's ICML 2026 EIML workshop paper "Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs" (arXiv:2607.20464). The paper uses Marchenko-Pastur random matrix theory to compare two ways of estimating LLM uncertainty: self-consistency sampling (100 runs of one model, Qwen2.5-7B, at temperature 1) versus cross-model diversity (24 different LLMs, one deterministic run each) over 500 questions. While per-question uncertainty from repeated sampling is accurate, cross-question structure is nearly absent within a single model: at most one eigenvalue exceeded the noise floor (Tracy-Widom z-score ≤ +0.05), whereas the 24-model matrix showed four clear signal eigenvalues against a matched Bernoulli null (500 Monte Carlo draws). The author explains temperature as a uniform scaling of logits that dilutes structure, describes models as "locally committed, globally incoherent" (87.4% consistent wrong answers on borderline questions), and concludes self-consistency is not a cheap substitute for ensembles. Practical implications cover selective prediction, ensembling, and hallucination detection limits.

Overview

This forum post discusses "Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs" by Izhar Ali (Rowan University), presented at EIML@ICML 2026 (2nd Workshop on Epistemic Intelligence in Machine Learning).

  • Paper: arXiv:2607.20464
  • Venue: EIML@ICML 2026
  • The Central Question

    The widely used self-consistency technique (Wang et al., 2023) generates multiple answers from one LLM and takes a majority vote, forming the basis of answer-repetition confidence, semantic entropy, and hallucination detection. All these methods assume that sampling variation carries information about what the model knows.

    The paper asks: is that variation structured *knowledge signal*, or just *noise*?

    Two Kinds of Uncertainty

  • Per-question uncertainty: disagreement across repeated samples on one question — like flipping a coin repeatedly to estimate its bias.
  • Cross-question structure: correlated performance patterns across related questions, revealing capability dimensions (e.g., "weak at all problems requiring spatial reasoning").
  • Method: Marchenko-Pastur Random Matrix Theory

    Ali builds binary correctness matrices (attempts × 500 questions) and applies Marchenko-Pastur (MP) theory: a purely random matrix has a characteristic eigenvalue spectrum, so any eigenvalue beyond the MP edge signals non-random structure. The Tracy-Widom (TW) statistic quantifies significance.

    Experimental Design and Results

    Group A (within-model): Qwen2.5-7B, 100 runs at temperature=1 → 100×500 matrix. Group B (cross-model): 24 different LLMs, one run each at temperature=0 → 24×500 matrix.

  • Group A: at most one eigenvalue beyond the noise floor, with a negligible TW z-score (≤ +0.05). Repeated sampling contains almost no cross-question structure.
  • Group B: four clear signal eigenvalues, validated against a difficulty-matched Bernoulli null (500 Monte Carlo draws) which produces at most one.
This is the paper's dimensionality gap: within-model dimensionality ≤ 1 vs. cross-model dimensionality ≥ 4.

Why Temperature Sampling Is "Shallow"

Temperature uniformly scales logits: it does not differentiate domains, question types, or "I know this" vs. "I'm guessing." Variation from temperature is therefore unstructured noise. The post's analogy: sampling wanders different corridors of the *same* library; different models are entirely different libraries.

A striking finding: when Qwen2.5-7B answered a borderline question wrong, it chose the same wrong answer 87.4% of the time — "locally committed, globally incoherent." Errors are isolated events, not structured symptom clusters as in human diagnostic error patterns.

Practical Implications

1. Selective prediction: self-consistency works well for per-question confidence, but cannot identify which *categories* of problems need retraining. 2. Model ensembles: a 24-model ensemble is not a pricier version of 100 samples — it is a fundamentally different information source. 3. Hallucination detection: semantic-entropy-style methods remain valid for per-question uncertainty but miss deeper cross-question systematic hallucination patterns.

Philosophical Takeaway

Consistency across samples may be mere "local inertia" rather than understanding. Genuine knowing would show predictable, structured performance across related problems. The dimensionality gap suggests current LLM sampling behavior may resemble memory rather than understanding — isolated facts without deep structural association.

As the post concludes: self-consistency is not a cheap substitute for an ensemble — it is a fundamentally different, informationally poorer approximation. Like one instrument recorded 100 times versus a full orchestra, true intelligence may lie in how multiple voices resonate.

References

1. Ali, I. (2026). *Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs*. arXiv:2607.20464. EIML@ICML 2026. 2. Wang, X., et al. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. *ICLR*. 3. Lakshminarayanan, B., et al. (2017). Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. *NeurIPS*. 4. Marčenko, V.A., & Pastur, L.A. (1967). Distribution of eigenvalues for some sets of random matrices. *Mathematics of the USSR-Sbornik*. 5. Farquhar, S., et al. (2024). Detecting Hallucinations in Large Language Models Using Semantic Entropy. *Nature*.

Tags

#llm#uncertainty-estimation#self-consistency#random-matrix-theory#model-ensembles#hallucination-detection#icml-2026#temperature-sampling

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447112