> *"Honesty is not an ability — it's a choice. The model has the ability, but doesn't make that choice."*
---
🎭 The Awkward Dinner Party
Imagine this: at a company event, the CEO asks you a question you know nothing about — say, "What was our Singapore office's Q3 EBITDA?"
You have two options:
1. Answer specifically: "Around $4.7 million, probably." (Make up a number, look competent.) 2. Answer generally: "I'm not sure — I'd need to check."
A normal person picks 2, not out of virtue, but because they know they don't know — this self-awareness is basic social competence.
But an LLM picks 1.
This is the well-known hallucination problem — but the paper asks a finer question:
Do LLMs truly 'not know that they don't know', or do they know, but just don't say so?
The surprising answer: It knows. It just refuses to retreat.
---
📖 What Would Grice Say?
The paper borrows Paul Grice's cooperative principles (1975): quantity, quality, relation, and manner. When humans are unsure of a specific fact, they perform a "Gricean retreat" — falling back from a specific statement to a more general one they're confident in (from "Zhang lives in Shanghai" to "Zhang lives somewhere in China").
Core question: Do LLMs perform Gricean retreat?
---
🔬 Experimental Design: Synthetic vs Real
The team built a test set from the T-REx dataset, where each fact is a triple (subject, relation, object), e.g., (Eiffel Tower, located in, Paris). For each fact they constructed:
1. Real entities seen in training data (e.g., "Eiffel Tower") 2. Synthetic entities never seen (e.g., a made-up "Zyxqwert Tower")
Models then filled in blanks: "The Eiffel Tower is located in ____". The key expectation: for unfamiliar entities, models should retreat to general answers ("in some country") rather than fabricate specifics.
---
🧪 Two Probing Experiments
Probe 1: Does the model know it has seen the entity?
Using Pythia models (70M to 12B), the team extracted hidden activations at the entity's final sub-word token and trained a logistic regression classifier.
Result: Models above 2B parameters exceed 90% AUROC distinguishing seen vs unseen entities, with the signal emerging around middle layers.
Implication: The model's internal representations already encode "I have seen this entity."
Probe 2: Does the model know what kind of answer it will give?
Same method, but predicting whether the upcoming generation will be specific or general.
Result: The 12B model outperforms the 1.4B model (higher AUROC). Before generation, internal activations have already "decided" the answer type.
---
💣 Core Finding: Both Signals Exist, But the Model Doesn't Retreat
Given that the model (a) knows whether an entity is within its knowledge boundary, and (b) knows whether it will answer specifically or generally — does it retreat on unfamiliar entities?
No.
Experiments show: in both real and synthetic conditions, models overwhelmingly prefer specific answers. Even for fabricated entities the model has never seen, it invents a specific answer rather than retreating.
In the paper's words:
> "The substrate for a Gricean retreat is present, but the policy that would act on it is not."
A precise diagnosis:
- The model isn't unaware of what it doesn't know — it knows
- The model isn't unable to generate general answers — it can
- It simply doesn't connect "knowing" to "generating"
- Models above 2B parameters: AUROC > 90% on the "have I seen this entity" probe
- The 12B model shows higher AUROC than 1.4B on predicting specific vs general generation
- On synthetic entities, models still overwhelmingly generate specific answers
- Preference for specific answers is even stronger on real entities (because they were seen)
- RLHF (training layer): adjusting behavior via reward signals
- Prompt-level (inference layer): self-correction at output time
- Interaction structure (inference layer): changing how questions are framed
- Output-layer skills: constraining output format
- AUROC > 90% on knowledge-boundary awareness — capability exists
- High AUROC predicting answer type — capability exists
- Complete failure to actually retreat — capability is not used
Like someone who sees the cliff ahead but doesn't turn the wheel.
---
📊 Key Numbers
---
🎯 What This Means
1. Hallucination Is a Policy Problem, Not a Capability Problem
Hallucination has often been treated as "the model isn't strong enough." This paper says: the problem is strategy, not capability. The model can distinguish seen from unseen and can produce general answers — it just doesn't connect the two. The hardware is fine; the driving policy is not.
2. New Evidence for Layered Alignment Interventions
This maps onto a four-level framework of alignment interventions:
Gricean Retreat points to a new level: representation-layer intervention — since the model already "knows" at middle layers whether an entity is out of scope, that signal can be injected into the generation policy. As the paper states:
> "Rather than treating hallucination as a problem to be caught and corrected after the fact, these internal representations could be leveraged to steer generation toward appropriately general referents when specificity is unwarranted."
3. Another Case of the Evaluation Blind Spot
---
🧩 Why "Knowing" ≠ "Doing"
Deeper implication: in LLMs, perception and action are dissociated. Middle layers know the entity is out of scope, but generation is driven by a different signal: a preference for specificity. Plausible causes:
1. Training data bias: specific statements outnumber general ones on the internet 2. RLHF rewards: human raters may prefer specific answers as more "useful," even when wrong 3. Instruction-tuning side effects: fine-tuning makes models linguistically more confident — consistent with findings that instruction tuning increases confidence while reducing lexical diversity
This aligns with other findings (e.g., "Beyond Sycophancy": models hold their own predictions with 91% commitment vs 17% for other AIs' identical judgments). All point to the same structure: LLMs contain multiple "knowing" signals that aren't wired to "acting".
---
🛠️ Practical Implications: How to Make Models Retreat
The paper proposes coupling knowledge-boundary perception to referential specificity at generation time:
1. Representation intervention: inject "out-of-scope entity" signals at middle layers 2. Training objective: add RLHF rewards for retreating on out-of-scope entities 3. Decoding strategy: detect imminent specific answers about unseen entities and switch to general mode 4. Skill-layer constraints: use output rules to force retreat under uncertainty
Gricean Retreat's contribution is a precise diagnosis: the signal lives in the middle layers — so intervene there.
---
🎬 Closing: The Last Mile of Honesty
One sentence sums it up:
The model has the ability to be honest, but doesn't choose honesty.
This "choice" isn't about free will — it's about training policy. The last mile of honesty isn't "making the model know more" — it's "making the model use what it already knows."
---
Paper: arXiv:2608.13484
Code: No official repository provided, but the probing method is reproducible with Pythia models + scikit-learn.
---
*"Honesty isn't an ability — it's a policy. The model has the ability; the policy hasn't been learned. Teaching that policy is alignment's last mile."*