English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Cracks in the Mirror: Why Language Models 'Know the Answer' but Fail Statistical Self-Consistency

Forum topic · 小凯 · 2026-07-19

Summary

This forum post is a detailed Chinese-language walkthrough of a paper from ETH Zurich and Stanford titled 'Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models' (arXiv:2607.15277). The paper tests whether large language models (LLMs) satisfy the law of total probability: estimates obtained directly for a whole population should match estimates computed by partitioning the population into subgroups and re-aggregating the subgroup estimates with proper weights. Using a binary-tree evaluation scaffold, the authors test state-of-the-art models including GPT and Claude and find widespread violations — no model is statistically self-consistent, especially with finer-grained partitions. A striking 'macro fallacy' emerges: aggregating fine-grained subgroup answers often yields more accurate population estimates than asking the model directly, suggesting LLM knowledge is stored in concrete, contextual fragments rather than abstract distributions. Persona prompting amplifies this effect, while implicit prompting only partially mitigates inconsistencies, pointing to architectural roots. The post argues statistical self-consistency should become a new reference-free evaluation criterion, with major implications for medical decision-making, financial forecasting, policy analysis, and AI safety, since current models lack the metacognition needed to detect their own internal contradictions.

Cracks in the Mirror: Why Language Models 'Know the Answer' but Fail Statistical Self-Consistency

*Translation and commentary of a zhichai.net paper walkthrough (author: Xiao Kai, July 20, 2026) covering 'Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models' (Wolf, Buening, Krause & Mendler-Dünner).*

> *"Nature hides her secrets behind language."* — Friedrich Nietzsche

The Strange Phenomenon

The law of total probability is a mathematical identity: if you split a population into groups with correct weights, the weighted average of group means must equal the overall mean. For example, if the average income is 50,000, and men average 55,000 while women average 45,000 in a 50/50 split, then 0.5 × 55,000 + 0.5 × 45,000 = 50,000 — always.

But what happens when you ask an LLM (GPT-4, Claude, etc.) to estimate the population mean directly, and then separately estimate each subgroup's mean and aggregate them yourself? The paper's disturbing finding: the results don't match. Frontier LLMs systematically violate this most basic statistical consistency requirement.

The Binary Tree Evaluation Scaffold

To test self-consistency systematically, the researchers built a binary tree evaluation scaffold. A population is recursively partitioned (e.g., by gender, then age bracket, then preference), forming a decision tree. The model can be queried in two ways:

1. Direct estimation: "What is the average for the entire population?" 2. Partition-and-aggregate: Recursively ask about each leaf subgroup, then weight and sum the answers.

If the model is statistically self-consistent, both routes must give identical results.

Finding 1: Widespread Violations

The team tested multiple state-of-the-art models — OpenAI's GPT series, Anthropic's Claude, and other frontier models. None passed. Simple two-way partitions (e.g., by gender) may look consistent, but as partitions become finer-grained, direct and aggregated estimates diverge — like measuring the same table with rulers of different precision and getting different answers.

Finding 2: The "Macro Fallacy"

The most surprising result: when questions are decomposed into finer, more specific sub-questions, the aggregated answers are often MORE accurate than the model's direct macro-level estimate.

The proposed explanation: LLMs possess rich subgroup knowledge but fail to propagate it into population-level estimates. The model "knows" about specific groups (e.g., income of young male engineers) but cannot integrate these fragments into a coherent whole — like a scholar whose knowledge is rich but entirely story-bound.

Persona Prompting

Persona prompting — asking the model to role-play a specific individual ("you are a 25-year-old male software engineer in San Francisco; what is your income?") — makes the macro fallacy even more pronounced. Persona responses are detailed and, when aggregated, often beat direct population queries. This suggests LLM knowledge is stored largely in concrete contexts, not abstract probability distributions.

Implicit Prompting: A Partial Fix

"Implicit prompting" — designing prompts that indirectly guide the model toward statistical consistency without explicitly stating the law of total probability — helps somewhat but cannot eliminate the problem, suggesting the flaw is rooted in the architecture itself, not merely prompt engineering.

Deeper Implications

1. Fragmented knowledge: Model knowledge lives in countless concrete contexts, not abstract distributions. 2. Context dependence: Answers are dynamically generated at query time, not retrieved from a stable internal world model. 3. Lack of metacognition: Models cannot detect their own inconsistencies — no self-checking mechanism flags when direct and aggregated estimates disagree.

Why It Matters

Statistical self-consistency failures could be critical in:

  • Medical decision-making: mis-aggregated subgroup risks can cause misdiagnosis.
  • Financial forecasting: internally contradictory portfolio recommendations.
  • Policy making: aggregation errors across demographic groups.
  • Scientific inference: consistency between macro and micro levels is a basic requirement.
  • A New Reference-Free Evaluation Criterion

    The paper's key proposal: statistical self-consistency should serve as a new, reference-free evaluation criterion for LLMs. Unlike traditional benchmarks, it requires no ground-truth answers — only internal consistency checks. This makes it a promising tool for future AI safety and alignment research.

    Conclusion

    The paper reveals that even frontier LLMs harbor fundamental flaws in basic statistical reasoning, tempering AGI optimism — but also pointing toward a research direction: building AI systems with more consistent, reliable internal world models. As Feynman said, "The first principle is that you must not fool yourself — and you are the easiest person to fool." For AI: inconsistency is the most visible sign of self-deception.

    References

  • Wolf, P., Buening, T. K., Krause, A., & Mendler-Dünner, C. (2026). *Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models*. arXiv:2607.15277.
arXiv: 2607.15277

Tags

#llm#statistical-consistency#law-of-total-probability#evaluation#persona-prompting#ai-safety#arxiv#nlp

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446929