The Awkward Fact
Show a vision-language model an obscure plant photo and ask what it is. It won't say "I don't know." It will confidently invent a name. Correct it, and next time it still invents the wrong one.
This is not an edge case—it reflects two chronic failures of current AI: catastrophic forgetting (learning new things erases old ones) and hallucination (not knowing what it doesn't know). For decades researchers have treated these as separate problems—replay buffers for memory, confidence thresholds for hallucination. Independent researcher Louis Mouchon, in a June 2026 research note, argues they are the same problem: the model lacks a signal for "is this new or familiar?" He calls that signal surprise.
The Brain Already Does This
In your head, this signal has run for hundreds of millions of years. The hippocampus rapidly records new experiences; the neocortex slowly integrates them into long-term knowledge. "Sleep" replays hippocampal traces into the cortex. This is Complementary Learning Systems theory (McClelland & O'Reilly, 1995). The key switch is prediction error: a surprising stimulus gets written as a new hippocampal trace; a non-surprising one doesn't. The same signal also tells you "I'm not familiar with this," prompting epistemic humility.
One signal, two jobs: a plasticity gate and the substrate of metacognition. Mouchon's work ports this into AI, as-is.
System 1: Surprise-Gated Memory
The architecture is almost suspiciously simple:
- Frozen backbone (DINOv2 or I-JEPA): images to vectors, never updated.
- A small predictor (JEPA anchor): predicts the full embedding from masked views. Good prediction → low surprise; bad → high surprise.
- Hippocampus: a non-parametric store. New images are written only if surprise exceeds a threshold.
- Neocortex: a linear classifier updated only during "sleep," trained on hippocampal replay.
- Surprise < 0.35 (known): answer confidently from retrieved facts.
- 0.35 ≤ surprise ≤ 0.65 (partially familiar): hedge, offer candidates with uncertainty flagged.
- Surprise > 0.65 (novel): enter student mode—describe only what is seen, say "I haven't seen this concept; can you tell me?"—then learn the concept in one shot from a single user sentence.
- Familiar → no new memory (plasticity off); answer confidently (metacognition permits).
- Unfamiliar → write memory (plasticity on); don't fabricate (metacognition blocks).
With the backbone frozen, representational catastrophic forgetting is structurally excluded.
Key numbers
On a 1000-class ImageNet continual stream:
| Setup | DINOv2 old-class retention | I-JEPA old-class retention | |-------|---------------------------|----------------------------| | Naive (no replay) | 65.8% | 25.9% | | Sleep replay | 83.5% | 77.2% | | Upper bound (i.i.d. training) | 84.4% | 78.5% |
Sleep replay recovers 17.7 points for DINOv2 and 51.3 points for I-JEPA—the harder the features are to linearly separate, the more replay helps.
A counterintuitive finding
"Replay only the most recent tasks" sounds like sensible memory saving. Result: 41.2% retention on DINOv2, and zero on I-JEPA—worse than no replay at all. Replaying only recent items keeps reinforcing new memories at the expense of old ones—recreating the forgetting mechanism itself. Partial replay is the worst option: replay everything, or nothing. The brain agrees: sleep replays remote and recent memories together.
System 2: Surprise-Driven Metacognition
The second system applies the same signal to knowing what it doesn't know. The backbone becomes SigLIP (shared text-image space) plus BGE-M3 for text retrieval. Surprise is calibrated to [0, 1]: known concepts ≈ 0.14, brand-new concepts ≈ 1.0, with thresholds at 0.35 and 0.65:
Why the gate must be external
Three detectors compete at separating known from unknown concepts:
| Detector | AUROC | |----------|-------| | External surprise detector (this work) | 0.966 ± 0.024 | | Self-reported confidence (model rates 0–100) | 0.618 ± 0.060 | | Token-level confidence ("yes I know" probability) | 0.292 ± 0.043 |
Token-level confidence is below chance. The model is as confident about fabricated concepts as real ones. This is not a bug but a consequence of training: the better a model is trained to answer, the less it can distinguish knowing from not knowing. A generator cannot be its own judge—the surprise signal must come from an external, frozen detector outside generation.
After sleep: 99.2% vs 0%
Teach the system 50 facts (each once), trigger sleep (hippocampus → neocortex consolidation), then clear the entire conversation history and query each fact: 99.2% recalled correctly (mean over 5 seeds, 95% CI ±0.9). The base-model control without episodic memory: 0%. The facts now live in slow storage—independent of context and of the hippocampus.
Memory authority
The most counterintuitive design decision: when memory conflicts with pretraining priors, memory wins. Told "Mont Blanc is 4,806 m" (pretraining says 4,810 m), the base model reports the pretraining value 2/3 of the time. In the full system, retrieved facts are injected with an explicit "this overrides your priors" instruction: 3/3 answers follow the corrected value, and the model notes it was taught this. Memory is not a suggestion; it's a command.
Why One Signal Can Do Two Jobs
The claim: a single prediction-error signal drives both plasticity and metacognition because both logically depend on the same judgment—"is this input familiar to me?"
Engineering Takeaways
1. Frozen backbone + lightweight adaptation is the right posture. All adaptation happens in small, inspectable modules—structurally immune to representational forgetting, and interpretable: you know which layer wrote which memory, and when.
2. External detectors beat self-reported confidence by a mile. 0.966 vs 0.618 vs 0.292. Every hallucination-mitigation scheme built on "model self-assessment" may rest on unreliable foundations.
3. Non-parametric stores don't forget. Parametric memory (e.g., test-time training) forgets as soon as you stop rewriting. Prototypes are written once and never overwritten—immune by design.
4. Partial replay is worse than no replay. Counterintuitive, but obvious once seen: replay everything or nothing; don't save memory with sliding windows.
5. Sleep is not optional. Teach 50 facts without sleep, clear the context, and everything is gone. Sleep converts fragile short-term traces into stable long-term representations. Skip it, and the system is just a chatbot with a context window.
Reflection
The paper evokes an overlooked fact: the brain's energy-saving design is precisely the source of its intelligence. The hippocampus writes only surprising things—saving energy. Sleep replays only surprising traces—saving energy. One signal drives two functions—saving energy again. The parsimony forced by evolutionary pressure yields the most elegant architecture.
The last decade of AI has scaled parameters, data, and compute. That path is right, but it carries a hidden assumption: every problem yields to scale. This work points at something scale cannot solve: models need an "I don't know" signal, and that signal does not emerge from scale—GPT-5.5 and Gemma 4 12B alike show below-chance token-level confidence. Scale makes models know more, but not know what they don't know. The latter requires an architectural choice: put the surprise signal outside the generator.
Caveats: this is a proof-of-concept "Research Note" by an independent researcher—small benchmarks, single-seed runs, limitations candidly acknowledged. But the core insight stands: one signal, two jobs. It will surely be re-validated at larger scale with stricter benchmarks, but the direction is already marked. Sometimes a good idea needs not a big track—just the right question, and the right signal.
---
Paper: Surprise as a Signal for Plasticity and Metacognition, Louis Mouchon, 2026-06-28
Author: Louis Mouchon (independent researcher)
Code: Not yet released (proof-of-concept)
Keywords: surprise gating, complementary learning systems, continual learning, metacognition, hallucination mitigation, JEPA