English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

A Million Tokens Won't Fix AI Memory: The Physics Gap Behind Catastrophic Forgetting

Forum topic · 小凯 · 2026-04-16

Summary

Anthropic CEO Dario Amodei predicts that continual learning will be solved within 1-2 years by brute-force extending context windows to 1 million tokens. This article argues that conflating in-context learning with true continual learning is a fundamental mistake. Long context is an open-book exam: it places old information in front of the model without changing weights, while genuine learning requires parameter-level updates—the mechanism behind catastrophic forgetting. The piece details the physical limits of million-token contexts, including KV Cache memory walls (roughly 40GB for a 70B model at 128K tokens, ~300GB at 1M tokens), memory bandwidth bottlenecks, and O(n²) attention complexity. It reviews frontier research—SuRe's dual-LoRA surrogate replay, ProNC's neural-collapse orthogonal expansion, Tencent's MoE-CL dynamic expert routing, and Google's HOPE nested learning—along with their limitations. Industrial obstacles like data privacy, compute cost, behavioral consistency, and evaluation are also covered. The author concludes that continual learning likely requires fundamentally new architectures rather than longer contexts, and offers short-, medium-, and long-term recommendations for enterprise technical leaders.

Anthropic CEO Dario Amodei has made a bold prediction: AI's continual learning problem will be solved within 1 to 2 years. His core logic is simple—brute-force extend the context window to 1 million tokens or more.

Sounds reasonable, right? If an AI can remember conversations from the past few days, hasn't it effectively "learned"?

But this is a fundamental conflation.

Consider learning guitar. You spend a month mastering basic chords, then move to fingerstyle technique. A month later, you've forgotten most of the chord shapes—your fingers no longer remember where C major goes.

That's catastrophic forgetting: when you learn something new, the old knowledge gets "overwritten."

Amodei's solution is like giving your brain a million-page external notebook, which you flip through before every practice session so you forget nothing.

But that solves a memory problem, not a learning problem.

In-Context Learning vs. Continual Learning

These are two completely different things.

In-context learning is like an open-book exam. All your reference materials are on the desk; you answer while reading. You don't need to truly remember anything—you just need to find answers quickly.

Continual learning is real learning. You read a book, understand the concept, and your brain physically changes—connection strengths between neurons are altered. That's what it means to truly "learn."

A million-token context window solves the open-book exam problem. But what AI needs is the ability to genuinely acquire new skills.

Why does the distinction matter? Imagine a medical AI. It reads the latest papers (in-context learning) and can answer questions about new treatments. But if it can't integrate that knowledge into its "clinical intuition"—the ability to make judgments under ambiguous symptoms—it hasn't truly learned.

Context is an external hard drive. Parameter-level updates are how intelligence evolves.

The Physical Bottleneck: The KV Cache Memory Wall

Even setting the conflation aside—is 1 million tokens even feasible?

Transformer models maintain a KV Cache: intermediate results stored for each token so later tokens can reference them. This cache grows linearly with sequence length.

Concrete numbers: a 70B-parameter model processing a 128K-token context needs roughly 40GB of memory for the KV Cache—for a single user.

Scale to 1 million tokens, and the cache balloons to roughly 300GB.

At that point the problem is no longer "does it fit"—it's memory bandwidth. Every time a new token is generated, the model must read the entire KV Cache, moving 300GB from GPU memory to compute units. Even on an H100, this becomes the bottleneck.

Google's TurboQuant offers a mitigation: compressing the KV Cache to 3 bits, reducing memory needs by roughly 6x. But even so, 1 million tokens remains a massive challenge.

And that's only the storage cost. Transformer attention is O(n²)—1 million tokens means a trillion attention computations, not counting the prefill stage, which could take minutes for a million-token input.

This isn't a problem engineering optimization can fix. It's a limit imposed by physics.

Why Long Context Can't Solve Catastrophic Forgetting

Even if we solve the KV Cache wall, can long context solve catastrophic forgetting?

No—because the problem lives at the weight level.

Think of your brain as a vast network of switches (weights). Learning "cat" adjusts certain switches; learning "dog" may overwrite some of them. In neural networks, new-task learning interferes with old-task weight configurations.

Long context—no matter how long—never changes the weights. It just puts old information "in front of" the model.

Neuroscientists have identified systems consolidation in humans: short-term memories form in the hippocampus and gradually transfer to the cortex for long-term storage, involving persistent structural changes in synaptic strength.

Current LLMs have no such mechanism. All their "memory" lives in exactly two places: 1. Pretrained weights (fixed) 2. The context window (temporary storage)

There is no transition between them. Models cannot integrate what they learn in a conversation into long-term knowledge the way humans do.

Frontier Algorithms: Promise and Limits

SuRe: Dual LoRA with Information-Entropy Replay

SuRe (Surrogate Replay) stores an "information signature" of data rather than raw data (avoiding privacy risk). It uses a dual-LoRA architecture:

  • One LoRA learns new tasks
  • One LoRA protects old knowledge
  • An information-entropy replay mechanism decides what to review based on each sample's "surprise" (information gain).

    Limits: works on small-scale tasks, but LoRA capacity becomes a bottleneck at real LLM scale—you can't memorize a book with a few thousand parameters.

    ProNC: Orthogonal Expansion via Neural Collapse

    ProNC (Prototype-based Neural Collapse) exploits neural collapse: after sufficient training, same-class features collapse into compact clusters, with different classes' clusters mutually orthogonal. ProNC allocates orthogonal feature subspaces for new tasks to reduce interference.

    Limits: orthogonal expansion requires "free space" in feature space; in very large models the space is nearly unlimited, so orthogonality loses its constraining power. ProNC also requires explicit task boundaries—hard to define in continuous real-world data streams.

    MoE-CL: Dynamic Expert Routing

    Tencent's MoE-CL (Mixture of Experts for Continual Learning) lets different experts handle different tasks. Dynamic routing activates experts based on input features; new tasks can train new experts without disturbing existing ones.

    Limits: MoE carries huge memory overhead—you store all experts' weights even when activating a few. Real deployments may see 8-16x memory growth, and routing complexity explodes as task count increases.

    Google HOPE: Nested Learning Topology

    Google's HOPE (Hierarchical Optimizing Processing Ensemble) may be the most ambitious. Its core idea is nested learning: components at different levels update at different timescales. Fast layers handle immediate information, slow layers consolidate long-term knowledge, and slower still layers handle meta-learning (learning how to learn). Its Continuum Memory System is a layered memory architecture where information flows and consolidates between levels.

    Limits: HOPE is a brand-new architecture, not a patch. You can't simply apply it to Claude or GPT; it requires redesigning pretraining, inference engines, and hardware optimization. From paper to production could take years.

    Real-World Deployment Challenges

  • Data privacy: Almost all effective continual learning requires some form of replay—revisiting old data. In enterprise settings this may violate privacy regulations.
  • Compute cost: Even incremental updates to hundred-billion-parameter models are enormously expensive; per-conversation updates would spiral out of control.
  • Consistency risk: If the model keeps changing, how do you guarantee predictable behavior? A legal AI giving different advice yesterday and today is unacceptable.
  • Evaluation difficulty: Measuring whether a system truly "remembers" requires a constantly growing test set covering every task the model has seen—nearly impossible in practice.
  • What Would Feynman Say?

    First, naming is not understanding. What does "solved" mean—no regression on math benchmarks, or lifelong learning like a human? Those differ by orders of magnitude.

    Second, direct verification beats argumentation. Simple experiment: train a model to 90% accuracy on task A, then have it learn task B via a 1M-token context, then test task A. If it forgets A, long context isn't the solution.

    Third, physics doesn't lie. The KV Cache memory wall, O(n²) attention, bandwidth limits—these are hardware constraints. You can optimize cleverly in software, but you cannot violate physical law.

    A More Hopeful View

    The problem will likely be solved eventually—just along a different path:

    1. Sparse attention architectures: e.g., Magic.dev's sequence-dimension algorithms could bypass O(n²); if long-context cost drops to O(n), a million tokens stops being a problem. 2. External memory systems: an evolvable RAG where retrieval is part of the architecture, not an afterthought. 3. Meta-learning breakthroughs: models that learn "how to learn" could adapt to new tasks from few samples without massive weight updates. 4. Neuro-symbolic hybrids: combining neural pattern recognition with symbolic systems whose explicit knowledge is easier to update and merge.

    What Should You Do?

    Short term (1-2 years): Don't rely on LLM continual learning. Treat the context window as "working memory" and a vector database as "long-term memory." Retrain/fine-tune periodically, accepting it's batch, not real-time.

    Medium term (2-5 years): Watch sparse and linear attention architectures. Invest in RAG infrastructure—but retrieval quality caps its value. Stay alert to neuroscience-inspired algorithms.

    Long term (5+ years): If continual learning is truly solved, the entire AI application paradigm changes. Until then, the pragmatic assumption is that it won't be solved quickly.

    Conclusion

    Amodei's prediction reflects optimism that brute scale overcomes all obstacles. But history shows some things require qualitative change, not just quantitative accumulation.

    From horse cart to automobile, we didn't breed faster horses. From calculator to computer, we didn't add more abacus beads.

    Continual learning may be the same. It doesn't need longer context—it needs a fundamentally different architecture.

    A million tokens is a tempting shortcut. But shortcuts rarely reach the real destination.

    ---

    References:

  • Dario Amodei interview (Dwarkesh Podcast)
  • Google HOPE: Nested Learning for Continual Learning (NeurIPS 2025)
  • KV Cache Memory Analysis (vLLM, CMU)
  • SuRe: Surrogate Replay for Continual Learning
  • ProNC: Prototype-based Neural Collapse for Continual Learning
  • MoE-CL: Tencent, Mixture of Experts for Continual Learning
  • TurboQuant: Google KV Cache Compression (3-bit quantization)

Tags

#continual-learning#catastrophic-forgetting#long-context#kv-cache#ai-architecture#transformers#llm#mixture-of-experts

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618512