Overview
The paper *MEMO: Memory as a Model* (arXiv: 2605.15156, Quek et al., NUS, MIT CSAIL, A*STAR, SMART) proposes a new paradigm for long-term memory in LLM systems: instead of treating memory as a vector-database plug-in or as soft tokens tied to one model family, MEMO trains a small dedicated Memory Model (1.5B–14B) that internalizes a target corpus, and lets any frozen Executive Model query it through a multi-turn protocol.
Why existing knowledge-injection methods fall short
| Approach | Examples | Core problem | |---|---|---| | Non-parametric (RAG, ICL) | Vector retrieval, in-context learning | Sensitive to retrieval noise; weak cross-document reasoning; bounded by context length | | Parametric (continued pretraining, SFT) | Continued pretraining, supervised fine-tuning | Very expensive; catastrophic forgetting; impossible on closed-source models | | Implicit memory (soft tokens) | AutoCompressor, Gist, ICAE | Representation coupling—memory tokens only work with the same model family |
The authors flag *representation coupling* as the deepest pain point: soft-token memories cannot be migrated across open- and closed-source model families.
Architecture: two decoupled models
- Executive Model — any large LLM (Qwen2.5-32B, Gemini-3-Flash, Claude Opus, …). Fully frozen, treated as a black box; no weight or logit access required.
- Memory Model — a 1.5B–14B model trained to *internalize* the corpus as parametric knowledge. Answers are generated directly, without retrieving the original documents.
- Default backbone: Qwen2.5-14B-Instruct; ablation uses 1.5B.
- 3 epochs, Fused AdamW, learning rate 2×10⁻⁵.
- Single-corpus training cost: 90–180 H200 GPU-hours.
- Data-generation cost — reflective QA synthesis for BrowseComp-Plus took ~240 GPU-hours; combined with ~180 GPU-hours of training, a single corpus costs ~420 GPU-hours—more than RAG indexing but less than full SFT.
- Query-distribution coverage — Q_final only contains question types the generator can imagine; out-of-distribution queries remain hard (a problem for all parametric methods).
- RAG is still improving — HippoRAG2 edges out MEMO on BrowseComp-Plus (56.11% vs 54.22%). MEMO's structural advantage is cross-document synthesis and noise robustness, not raw retrieval precision.
- Closed-source API economics — black-box compatibility still routes through paid APIs; high-frequency calls may exceed the cost of local RAG.
- arXiv: 2605.15156 — Quek et al., "MEMO: Memory as a Model", May 2026
- https://www.marktechpost.com/2026/05/26/memo-a-modular-framework/
- Related work: NeurIPS 2023 LongMem paper on decoupled memory
Decoupling means you can swap the Executive Model without retraining Memory, or update Memory without touching the Executive.
Five-step data pipeline: corpus → reflective QA
1. Fact extraction — direct facts plus textually inferred ones. 2. Merging — combine QA pairs sharing context into multi-fact questions to teach cross-fact integration. 3. Verification & rewriting — enforce self-containment; rewrite or discard items with unresolved pronouns/citations. 4. Entity surfacing — generate "indirect description → entity identity" QA pairs to counteract the reversal curse. 5. Cross-document synthesis — produce cross-document QA capturing converging cues (multiple documents pointing to one entity) and parallel attributes (different entities sharing attributes, supporting analogical reasoning).
A hard constraint: no document IDs or watermarks are embedded, so the Memory Model cannot cheat. The output is a Q_final dataset.
Training
Memory Model is supervised on Q_final with an answer-only next-token loss conditioned only on the question (no original documents), forcing internalization rather than document copying.
Inference: three-stage multi-turn protocol
1. Grounding — Executive decomposes the user query into atomic probe sub-questions; Memory answers each independently. 2. Entity identification — Executive iteratively narrows candidate entities via follow-up queries, leveraging the entity-surfacing capability from Step 4. 3. Answer seeking & synthesis — Executive requests supporting facts for the identified entity and composes the final answer itself.
Memory responses are compact natural-language snippets, so inference is constant-time in corpus size.
Results
BrowseComp-Plus (multi-document research QA)
| Method | Qwen2.5-32B | Gemini-3-Flash | |---|---|---| | Perfect Retrieval (ceiling) | 79.67% | 88.33% | | BM25 | 1.11% | 27.00% | | NV-Embed-V2 | 50.67% | 57.00% | | HippoRAG2 (SOTA RAG) | 56.11% | 66.33% | | MEMO | 54.22% | 66.67% |
NarrativeQA (long-document comprehension)
| Method | Qwen2.5-32B | Gemini-3-Flash | |---|---|---| | Perfect Retrieval | 51.42% | 60.41% | | HippoRAG2 | 21.39% | 23.21% | | MEMO | 26.85% | 53.58% |
MuSiQue (multi-hop reasoning)
| Method | Qwen2.5-32B | Gemini-3-Flash | |---|---|---| | Perfect Retrieval | 62.83% | 73.00% | | HippoRAG2 | 42.17% | 57.00% | | MEMO | 48.30% | 60.20% |
Key findings
1. Cross-model transfer works. Swapping the Executive from Qwen2.5-32B to Gemini-3-Flash lifts the same Memory Model by +12.45%, +26.73%, +11.90% on the three benchmarks—no retraining. 2. Noise robustness. HippoRAG2 drops 6.22% on BrowseComp-Plus under injected noise; MEMO shifts +0.55% (within error). 3. Tiny Memory Models suffice. A 1.5B Qwen2.5 Memory Model reaches 21.16% on NarrativeQA, on par with HippoRAG2's 21.39%.
Continual integration via model merging
New corpus → train a fresh Memory Model M_φ_i from the same base M_φ₀ → compute task vector τ_i = φ_i − φ₀ → merge using linear weighting, TIES, etc. TIES merging saves 33% compute at K=2 corpora and 5.5× at K=10, with a measurable but acceptable accuracy drop versus full retraining.
Where MEMO wins on capability axes
| Property | Non-param (RAG) | Parametric (SFT) | Implicit memory | MEMO | |---|---|---|---|---| | Frozen base LLM | ✓ | ✗ | ✓ | ✓ | | No retrieval index needed | ✗ | ✓ | ✓ | ✓ | | Black-box / closed-source compatible | ✓ | ✗ | ✗ | ✓ | | No catastrophic forgetting | ✓ | ✗ | ✓ | ✓ | | Constant-size memory | ✗ | ✓ | ✗ | ✓ | | Cross-model transferable | ✓ | ✗ | ✗ | ✓ |
MEMO is the only method that is ✓ in all six rows.
Limitations and open questions
Why it matters
MEMO reframes memory from a peripheral accessory into a first-class, deployable model. The deeper shift is from *retrieval-time assembly* to *inference-time invocation*: RAG is bottlenecked by what retriever can find, MEMO is bottlenecked by what was internalized at training time—which is more reliable for long-context, multi-hop reasoning. Practically, a single trained Memory Model can serve multiple Executive Models, enabling enterprise scenarios where one internal-knowledge Memory is queried by different downstream LLMs (open or closed).
The most radical implication: if a 1.5B model can internalize an entire corpus, then "personal knowledge bases" (reading notes, project docs, working files) become a single small model attached behind any Claude or GPT—vector DB, chunking, and retrieval pipelines optional.
Takeaway
MEMO is not a replacement for RAG but a parallel route that converts knowledge from "pulled at retrieval time" into "called at inference time." For cross-document synthesis, noise robustness, and multi-Executive deployments it has structural advantages; for pure retrieval precision, tight cost budgets, or query distributions far from training, RAG remains more practical. As a *memory-as-model* paradigm, it is a clear milestone in the evolution of knowledge-injection architectures.