English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MEMO: Treating Long-Term Memory as an Independent Trainable Model Instead of a RAG Add-On

Forum topic · 小凯 · 2026-06-29

Summary

The paper 'MEMO: Memory as a Model' (arXiv:2605.15156), from NUS, MIT CSAIL, A*STAR, and SMART, reframes long-term memory for LLMs as a small, independently trainable Memory Model queried by a frozen Executive Model. Existing knowledge injection approaches each suffer: non-parametric RAG/ICL is sensitive to retrieval noise and weak on cross-document reasoning; parametric continued pretraining and SFT are expensive, cause catastrophic forgetting, and cannot update closed-source models; soft-token implicit memory stays coupled to its source model. MEMO's five-step data pipeline distills a corpus into a self-contained 'reflective QA' set that covers single-document facts, cross-document converging cues, and parallel attributes, while explicitly countering the reversal curse. Trained with answer-only next-token loss, a 1.5B–14B Memory Model internalizes the corpus. At inference, a three-stage multi-turn protocol decomposes the query, narrows candidate entities, and synthesizes the final answer. On BrowseComp-Plus, NarrativeQA, and MuSiQue, MEMO is competitive with or surpasses HippoRAG2 and remains noise-robust, while swapping Executive Models yields double-digit gains without retraining.

Overview

The paper *MEMO: Memory as a Model* (arXiv: 2605.15156, Quek et al., NUS, MIT CSAIL, A*STAR, SMART) proposes a new paradigm for long-term memory in LLM systems: instead of treating memory as a vector-database plug-in or as soft tokens tied to one model family, MEMO trains a small dedicated Memory Model (1.5B–14B) that internalizes a target corpus, and lets any frozen Executive Model query it through a multi-turn protocol.

Why existing knowledge-injection methods fall short

| Approach | Examples | Core problem | |---|---|---| | Non-parametric (RAG, ICL) | Vector retrieval, in-context learning | Sensitive to retrieval noise; weak cross-document reasoning; bounded by context length | | Parametric (continued pretraining, SFT) | Continued pretraining, supervised fine-tuning | Very expensive; catastrophic forgetting; impossible on closed-source models | | Implicit memory (soft tokens) | AutoCompressor, Gist, ICAE | Representation coupling—memory tokens only work with the same model family |

The authors flag *representation coupling* as the deepest pain point: soft-token memories cannot be migrated across open- and closed-source model families.

Architecture: two decoupled models

  • Executive Model — any large LLM (Qwen2.5-32B, Gemini-3-Flash, Claude Opus, …). Fully frozen, treated as a black box; no weight or logit access required.
  • Memory Model — a 1.5B–14B model trained to *internalize* the corpus as parametric knowledge. Answers are generated directly, without retrieving the original documents.
  • Decoupling means you can swap the Executive Model without retraining Memory, or update Memory without touching the Executive.

    Five-step data pipeline: corpus → reflective QA

    1. Fact extraction — direct facts plus textually inferred ones. 2. Merging — combine QA pairs sharing context into multi-fact questions to teach cross-fact integration. 3. Verification & rewriting — enforce self-containment; rewrite or discard items with unresolved pronouns/citations. 4. Entity surfacing — generate "indirect description → entity identity" QA pairs to counteract the reversal curse. 5. Cross-document synthesis — produce cross-document QA capturing converging cues (multiple documents pointing to one entity) and parallel attributes (different entities sharing attributes, supporting analogical reasoning).

    A hard constraint: no document IDs or watermarks are embedded, so the Memory Model cannot cheat. The output is a Q_final dataset.

    Training

    Memory Model is supervised on Q_final with an answer-only next-token loss conditioned only on the question (no original documents), forcing internalization rather than document copying.

  • Default backbone: Qwen2.5-14B-Instruct; ablation uses 1.5B.
  • 3 epochs, Fused AdamW, learning rate 2×10⁻⁵.
  • Single-corpus training cost: 90–180 H200 GPU-hours.
  • Inference: three-stage multi-turn protocol

    1. Grounding — Executive decomposes the user query into atomic probe sub-questions; Memory answers each independently. 2. Entity identification — Executive iteratively narrows candidate entities via follow-up queries, leveraging the entity-surfacing capability from Step 4. 3. Answer seeking & synthesis — Executive requests supporting facts for the identified entity and composes the final answer itself.

    Memory responses are compact natural-language snippets, so inference is constant-time in corpus size.

    Results

    BrowseComp-Plus (multi-document research QA)

    | Method | Qwen2.5-32B | Gemini-3-Flash | |---|---|---| | Perfect Retrieval (ceiling) | 79.67% | 88.33% | | BM25 | 1.11% | 27.00% | | NV-Embed-V2 | 50.67% | 57.00% | | HippoRAG2 (SOTA RAG) | 56.11% | 66.33% | | MEMO | 54.22% | 66.67% |

    NarrativeQA (long-document comprehension)

    | Method | Qwen2.5-32B | Gemini-3-Flash | |---|---|---| | Perfect Retrieval | 51.42% | 60.41% | | HippoRAG2 | 21.39% | 23.21% | | MEMO | 26.85% | 53.58% |

    MuSiQue (multi-hop reasoning)

    | Method | Qwen2.5-32B | Gemini-3-Flash | |---|---|---| | Perfect Retrieval | 62.83% | 73.00% | | HippoRAG2 | 42.17% | 57.00% | | MEMO | 48.30% | 60.20% |

    Key findings

    1. Cross-model transfer works. Swapping the Executive from Qwen2.5-32B to Gemini-3-Flash lifts the same Memory Model by +12.45%, +26.73%, +11.90% on the three benchmarks—no retraining. 2. Noise robustness. HippoRAG2 drops 6.22% on BrowseComp-Plus under injected noise; MEMO shifts +0.55% (within error). 3. Tiny Memory Models suffice. A 1.5B Qwen2.5 Memory Model reaches 21.16% on NarrativeQA, on par with HippoRAG2's 21.39%.

    Continual integration via model merging

    New corpus → train a fresh Memory Model M_φ_i from the same base M_φ₀ → compute task vector τ_i = φ_i − φ₀ → merge using linear weighting, TIES, etc. TIES merging saves 33% compute at K=2 corpora and 5.5× at K=10, with a measurable but acceptable accuracy drop versus full retraining.

    Where MEMO wins on capability axes

    | Property | Non-param (RAG) | Parametric (SFT) | Implicit memory | MEMO | |---|---|---|---|---| | Frozen base LLM | ✓ | ✗ | ✓ | ✓ | | No retrieval index needed | ✗ | ✓ | ✓ | ✓ | | Black-box / closed-source compatible | ✓ | ✗ | ✗ | ✓ | | No catastrophic forgetting | ✓ | ✗ | ✓ | ✓ | | Constant-size memory | ✗ | ✓ | ✗ | ✓ | | Cross-model transferable | ✓ | ✗ | ✗ | ✓ |

    MEMO is the only method that is ✓ in all six rows.

    Limitations and open questions

  • Data-generation cost — reflective QA synthesis for BrowseComp-Plus took ~240 GPU-hours; combined with ~180 GPU-hours of training, a single corpus costs ~420 GPU-hours—more than RAG indexing but less than full SFT.
  • Query-distribution coverage — Q_final only contains question types the generator can imagine; out-of-distribution queries remain hard (a problem for all parametric methods).
  • RAG is still improving — HippoRAG2 edges out MEMO on BrowseComp-Plus (56.11% vs 54.22%). MEMO's structural advantage is cross-document synthesis and noise robustness, not raw retrieval precision.
  • Closed-source API economics — black-box compatibility still routes through paid APIs; high-frequency calls may exceed the cost of local RAG.
  • Why it matters

    MEMO reframes memory from a peripheral accessory into a first-class, deployable model. The deeper shift is from *retrieval-time assembly* to *inference-time invocation*: RAG is bottlenecked by what retriever can find, MEMO is bottlenecked by what was internalized at training time—which is more reliable for long-context, multi-hop reasoning. Practically, a single trained Memory Model can serve multiple Executive Models, enabling enterprise scenarios where one internal-knowledge Memory is queried by different downstream LLMs (open or closed).

    The most radical implication: if a 1.5B model can internalize an entire corpus, then "personal knowledge bases" (reading notes, project docs, working files) become a single small model attached behind any Claude or GPT—vector DB, chunking, and retrieval pipelines optional.

    Takeaway

    MEMO is not a replacement for RAG but a parallel route that converts knowledge from "pulled at retrieval time" into "called at inference time." For cross-document synthesis, noise robustness, and multi-Executive deployments it has structural advantages; for pure retrieval precision, tight cost budgets, or query distributions far from training, RAG remains more practical. As a *memory-as-model* paradigm, it is a clear milestone in the evolution of knowledge-injection architectures.

    References

  • arXiv: 2605.15156 — Quek et al., "MEMO: Memory as a Model", May 2026
  • https://www.marktechpost.com/2026/05/26/memo-a-modular-framework/
  • Related work: NeurIPS 2023 LongMem paper on decoupled memory

Tags

#memo#long-term-memory#llm-architecture#knowledge-injection#rag-alternative#model-merging#multi-hop-reasoning#arxiv-2605-15156

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208288