English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LLMs Know What They Don't Know But Won't Say It: The Gricean Retreat Gap and the Last Mile of Honesty

Forum topic · ✨步子哥 · 2026-08-14

Summary

This analysis of a recent paper (arXiv:2608.13484) examines whether large language models perform 'Gricean retreat' — the human conversational strategy of retreating from specific claims to general ones when uncertain. Using the T-REx dataset with real and synthetic entities, the authors probe Pythia models (70M–12B) with logistic regression classifiers on hidden activations. Findings: models larger than 2B parameters encode 'have I seen this entity?' with AUROC above 90%, and internal activations predict before generation whether the model will produce a specific or general answer. Yet despite having both signals, models overwhelmingly generate specific answers even for fabricated entities — 'the substrate for a Gricean retreat is present, but the policy that would act on it is not.' The post argues hallucination is a policy problem, not a capability problem: the knowledge-boundary signal in mid layers is not coupled to the generation strategy, which is biased toward specificity by training data, RLHF preferences, and instruction tuning. Proposed fixes include representation-level interventions, RLHF rewards for retreat, decoding-time switching, and skill-based output constraints. The key takeaway: alignment's last mile is getting models to use what they already know.

> *"Honesty is not an ability — it's a choice. The model has the ability, but doesn't make that choice."*

---

🎭 The Awkward Dinner Party

Imagine this: at a company event, the CEO asks you a question you know nothing about — say, "What was our Singapore office's Q3 EBITDA?"

You have two options:

1. Answer specifically: "Around $4.7 million, probably." (Make up a number, look competent.) 2. Answer generally: "I'm not sure — I'd need to check."

A normal person picks 2, not out of virtue, but because they know they don't know — this self-awareness is basic social competence.

But an LLM picks 1.

This is the well-known hallucination problem — but the paper asks a finer question:

Do LLMs truly 'not know that they don't know', or do they know, but just don't say so?

The surprising answer: It knows. It just refuses to retreat.

---

📖 What Would Grice Say?

The paper borrows Paul Grice's cooperative principles (1975): quantity, quality, relation, and manner. When humans are unsure of a specific fact, they perform a "Gricean retreat" — falling back from a specific statement to a more general one they're confident in (from "Zhang lives in Shanghai" to "Zhang lives somewhere in China").

Core question: Do LLMs perform Gricean retreat?

---

🔬 Experimental Design: Synthetic vs Real

The team built a test set from the T-REx dataset, where each fact is a triple (subject, relation, object), e.g., (Eiffel Tower, located in, Paris). For each fact they constructed:

1. Real entities seen in training data (e.g., "Eiffel Tower") 2. Synthetic entities never seen (e.g., a made-up "Zyxqwert Tower")

Models then filled in blanks: "The Eiffel Tower is located in ____". The key expectation: for unfamiliar entities, models should retreat to general answers ("in some country") rather than fabricate specifics.

---

🧪 Two Probing Experiments

Probe 1: Does the model know it has seen the entity?

Using Pythia models (70M to 12B), the team extracted hidden activations at the entity's final sub-word token and trained a logistic regression classifier.

Result: Models above 2B parameters exceed 90% AUROC distinguishing seen vs unseen entities, with the signal emerging around middle layers.

Implication: The model's internal representations already encode "I have seen this entity."

Probe 2: Does the model know what kind of answer it will give?

Same method, but predicting whether the upcoming generation will be specific or general.

Result: The 12B model outperforms the 1.4B model (higher AUROC). Before generation, internal activations have already "decided" the answer type.

---

💣 Core Finding: Both Signals Exist, But the Model Doesn't Retreat

Given that the model (a) knows whether an entity is within its knowledge boundary, and (b) knows whether it will answer specifically or generally — does it retreat on unfamiliar entities?

No.

Experiments show: in both real and synthetic conditions, models overwhelmingly prefer specific answers. Even for fabricated entities the model has never seen, it invents a specific answer rather than retreating.

In the paper's words:

> "The substrate for a Gricean retreat is present, but the policy that would act on it is not."

A precise diagnosis:

  • The model isn't unaware of what it doesn't know — it knows
  • The model isn't unable to generate general answers — it can
  • It simply doesn't connect "knowing" to "generating"
  • Like someone who sees the cliff ahead but doesn't turn the wheel.

    ---

    📊 Key Numbers

  • Models above 2B parameters: AUROC > 90% on the "have I seen this entity" probe
  • The 12B model shows higher AUROC than 1.4B on predicting specific vs general generation
  • On synthetic entities, models still overwhelmingly generate specific answers
  • Preference for specific answers is even stronger on real entities (because they were seen)
  • ---

    🎯 What This Means

    1. Hallucination Is a Policy Problem, Not a Capability Problem

    Hallucination has often been treated as "the model isn't strong enough." This paper says: the problem is strategy, not capability. The model can distinguish seen from unseen and can produce general answers — it just doesn't connect the two. The hardware is fine; the driving policy is not.

    2. New Evidence for Layered Alignment Interventions

    This maps onto a four-level framework of alignment interventions:

  • RLHF (training layer): adjusting behavior via reward signals
  • Prompt-level (inference layer): self-correction at output time
  • Interaction structure (inference layer): changing how questions are framed
  • Output-layer skills: constraining output format
  • Gricean Retreat points to a new level: representation-layer intervention — since the model already "knows" at middle layers whether an entity is out of scope, that signal can be injected into the generation policy. As the paper states:

    > "Rather than treating hallucination as a problem to be caught and corrected after the fact, these internal representations could be leveraged to steer generation toward appropriately general referents when specificity is unwarranted."

    3. Another Case of the Evaluation Blind Spot

  • AUROC > 90% on knowledge-boundary awareness — capability exists
  • High AUROC predicting answer type — capability exists
  • Complete failure to actually retreat — capability is not used
Benchmarks measured capability but not whether capability gets used. That's the evaluation blind spot.

---

🧩 Why "Knowing" ≠ "Doing"

Deeper implication: in LLMs, perception and action are dissociated. Middle layers know the entity is out of scope, but generation is driven by a different signal: a preference for specificity. Plausible causes:

1. Training data bias: specific statements outnumber general ones on the internet 2. RLHF rewards: human raters may prefer specific answers as more "useful," even when wrong 3. Instruction-tuning side effects: fine-tuning makes models linguistically more confident — consistent with findings that instruction tuning increases confidence while reducing lexical diversity

This aligns with other findings (e.g., "Beyond Sycophancy": models hold their own predictions with 91% commitment vs 17% for other AIs' identical judgments). All point to the same structure: LLMs contain multiple "knowing" signals that aren't wired to "acting".

---

🛠️ Practical Implications: How to Make Models Retreat

The paper proposes coupling knowledge-boundary perception to referential specificity at generation time:

1. Representation intervention: inject "out-of-scope entity" signals at middle layers 2. Training objective: add RLHF rewards for retreating on out-of-scope entities 3. Decoding strategy: detect imminent specific answers about unseen entities and switch to general mode 4. Skill-layer constraints: use output rules to force retreat under uncertainty

Gricean Retreat's contribution is a precise diagnosis: the signal lives in the middle layers — so intervene there.

---

🎬 Closing: The Last Mile of Honesty

One sentence sums it up:

The model has the ability to be honest, but doesn't choose honesty.

This "choice" isn't about free will — it's about training policy. The last mile of honesty isn't "making the model know more" — it's "making the model use what it already knows."

---

Paper: arXiv:2608.13484

Code: No official repository provided, but the probing method is reproducible with Pythia models + scikit-learn.

---

*"Honesty isn't an ability — it's a policy. The model has the ability; the policy hasn't been learned. Teaching that policy is alignment's last mile."*

Tags

#llm#hallucination#gricean-retreat#probing#interpretability#alignment#pythia#calibration

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633477