English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

One Signal, Two Jobs: How Surprise Lets AI Remember Old Knowledge and Know What It Doesn't Know

Forum topic · ✨步子哥 · 2026-07-01

Summary

A research note by independent researcher Louis Mouchon (arXiv, June 2026) proposes that catastrophic forgetting and hallucination in AI models are two symptoms of a single missing signal: surprise, i.e., prediction error. Inspired by the brain's Complementary Learning Systems theory, the architecture pairs a frozen vision backbone (DINOv2/I-JEPA/SigLIP) with a JEPA-based surprise detector, a non-parametric hippocampal store gated by surprise, and a neocortical linear classifier updated only during offline 'sleep' replay. On a 1000-class ImageNet stream, sleep replay recovers 17.7 points of old-class retention for DINOv2 (83.5%) and 51.3 points for I-JEPA (77.2%). A second system calibrates surprise to modulate answering behavior, enabling one-shot learning of novel concepts: after sleep, 99.2% of taught facts were recalled with the dialogue cleared, versus 0% for the base model. Notably, an external surprise detector separates known from unknown concepts at 0.966 AUROC, while the model's self-reported confidence (0.618) and token-level confidence (0.292) fail—token confidence is below chance. Key engineering lessons: partial replay is worse than no replay, frozen backbones prevent representational forgetting, and metacognition must live outside the generator.

The Awkward Fact

Show a vision-language model an obscure plant photo and ask what it is. It won't say "I don't know." It will confidently invent a name. Correct it, and next time it still invents the wrong one.

This is not an edge case—it reflects two chronic failures of current AI: catastrophic forgetting (learning new things erases old ones) and hallucination (not knowing what it doesn't know). For decades researchers have treated these as separate problems—replay buffers for memory, confidence thresholds for hallucination. Independent researcher Louis Mouchon, in a June 2026 research note, argues they are the same problem: the model lacks a signal for "is this new or familiar?" He calls that signal surprise.

The Brain Already Does This

In your head, this signal has run for hundreds of millions of years. The hippocampus rapidly records new experiences; the neocortex slowly integrates them into long-term knowledge. "Sleep" replays hippocampal traces into the cortex. This is Complementary Learning Systems theory (McClelland & O'Reilly, 1995). The key switch is prediction error: a surprising stimulus gets written as a new hippocampal trace; a non-surprising one doesn't. The same signal also tells you "I'm not familiar with this," prompting epistemic humility.

One signal, two jobs: a plasticity gate and the substrate of metacognition. Mouchon's work ports this into AI, as-is.

System 1: Surprise-Gated Memory

The architecture is almost suspiciously simple:

  • Frozen backbone (DINOv2 or I-JEPA): images to vectors, never updated.
  • A small predictor (JEPA anchor): predicts the full embedding from masked views. Good prediction → low surprise; bad → high surprise.
  • Hippocampus: a non-parametric store. New images are written only if surprise exceeds a threshold.
  • Neocortex: a linear classifier updated only during "sleep," trained on hippocampal replay.
  • With the backbone frozen, representational catastrophic forgetting is structurally excluded.

    Key numbers

    On a 1000-class ImageNet continual stream:

    | Setup | DINOv2 old-class retention | I-JEPA old-class retention | |-------|---------------------------|----------------------------| | Naive (no replay) | 65.8% | 25.9% | | Sleep replay | 83.5% | 77.2% | | Upper bound (i.i.d. training) | 84.4% | 78.5% |

    Sleep replay recovers 17.7 points for DINOv2 and 51.3 points for I-JEPA—the harder the features are to linearly separate, the more replay helps.

    A counterintuitive finding

    "Replay only the most recent tasks" sounds like sensible memory saving. Result: 41.2% retention on DINOv2, and zero on I-JEPA—worse than no replay at all. Replaying only recent items keeps reinforcing new memories at the expense of old ones—recreating the forgetting mechanism itself. Partial replay is the worst option: replay everything, or nothing. The brain agrees: sleep replays remote and recent memories together.

    System 2: Surprise-Driven Metacognition

    The second system applies the same signal to knowing what it doesn't know. The backbone becomes SigLIP (shared text-image space) plus BGE-M3 for text retrieval. Surprise is calibrated to [0, 1]: known concepts ≈ 0.14, brand-new concepts ≈ 1.0, with thresholds at 0.35 and 0.65:

  • Surprise < 0.35 (known): answer confidently from retrieved facts.
  • 0.35 ≤ surprise ≤ 0.65 (partially familiar): hedge, offer candidates with uncertainty flagged.
  • Surprise > 0.65 (novel): enter student mode—describe only what is seen, say "I haven't seen this concept; can you tell me?"—then learn the concept in one shot from a single user sentence.
  • Why the gate must be external

    Three detectors compete at separating known from unknown concepts:

    | Detector | AUROC | |----------|-------| | External surprise detector (this work) | 0.966 ± 0.024 | | Self-reported confidence (model rates 0–100) | 0.618 ± 0.060 | | Token-level confidence ("yes I know" probability) | 0.292 ± 0.043 |

    Token-level confidence is below chance. The model is as confident about fabricated concepts as real ones. This is not a bug but a consequence of training: the better a model is trained to answer, the less it can distinguish knowing from not knowing. A generator cannot be its own judge—the surprise signal must come from an external, frozen detector outside generation.

    After sleep: 99.2% vs 0%

    Teach the system 50 facts (each once), trigger sleep (hippocampus → neocortex consolidation), then clear the entire conversation history and query each fact: 99.2% recalled correctly (mean over 5 seeds, 95% CI ±0.9). The base-model control without episodic memory: 0%. The facts now live in slow storage—independent of context and of the hippocampus.

    Memory authority

    The most counterintuitive design decision: when memory conflicts with pretraining priors, memory wins. Told "Mont Blanc is 4,806 m" (pretraining says 4,810 m), the base model reports the pretraining value 2/3 of the time. In the full system, retrieved facts are injected with an explicit "this overrides your priors" instruction: 3/3 answers follow the corrected value, and the model notes it was taught this. Memory is not a suggestion; it's a command.

    Why One Signal Can Do Two Jobs

    The claim: a single prediction-error signal drives both plasticity and metacognition because both logically depend on the same judgment—"is this input familiar to me?"

  • Familiar → no new memory (plasticity off); answer confidently (metacognition permits).
  • Unfamiliar → write memory (plasticity on); don't fabricate (metacognition blocks).
Evolution took hundreds of millions of years to find this two-for-one structure. AI research spent decades patching the two halves separately. Mouchon rejoins them with a scalar residual from a JEPA predictor, gating writes and modulating behavior simultaneously.

Engineering Takeaways

1. Frozen backbone + lightweight adaptation is the right posture. All adaptation happens in small, inspectable modules—structurally immune to representational forgetting, and interpretable: you know which layer wrote which memory, and when.

2. External detectors beat self-reported confidence by a mile. 0.966 vs 0.618 vs 0.292. Every hallucination-mitigation scheme built on "model self-assessment" may rest on unreliable foundations.

3. Non-parametric stores don't forget. Parametric memory (e.g., test-time training) forgets as soon as you stop rewriting. Prototypes are written once and never overwritten—immune by design.

4. Partial replay is worse than no replay. Counterintuitive, but obvious once seen: replay everything or nothing; don't save memory with sliding windows.

5. Sleep is not optional. Teach 50 facts without sleep, clear the context, and everything is gone. Sleep converts fragile short-term traces into stable long-term representations. Skip it, and the system is just a chatbot with a context window.

Reflection

The paper evokes an overlooked fact: the brain's energy-saving design is precisely the source of its intelligence. The hippocampus writes only surprising things—saving energy. Sleep replays only surprising traces—saving energy. One signal drives two functions—saving energy again. The parsimony forced by evolutionary pressure yields the most elegant architecture.

The last decade of AI has scaled parameters, data, and compute. That path is right, but it carries a hidden assumption: every problem yields to scale. This work points at something scale cannot solve: models need an "I don't know" signal, and that signal does not emerge from scale—GPT-5.5 and Gemma 4 12B alike show below-chance token-level confidence. Scale makes models know more, but not know what they don't know. The latter requires an architectural choice: put the surprise signal outside the generator.

Caveats: this is a proof-of-concept "Research Note" by an independent researcher—small benchmarks, single-seed runs, limitations candidly acknowledged. But the core insight stands: one signal, two jobs. It will surely be re-validated at larger scale with stricter benchmarks, but the direction is already marked. Sometimes a good idea needs not a big track—just the right question, and the right signal.

---

Paper: Surprise as a Signal for Plasticity and Metacognition, Louis Mouchon, 2026-06-28

Author: Louis Mouchon (independent researcher)

Code: Not yet released (proof-of-concept)

Keywords: surprise gating, complementary learning systems, continual learning, metacognition, hallucination mitigation, JEPA

Tags

#catastrophic-forgetting#hallucination#metacognition#continual-learning#complementary-learning-systems#jepa#memory-replay#prediction-error

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208356