I came across a "Daily Arxiv" video on Bilibili: AI Finally Says Goodbye to Goldfish Memory! Metis: A Native Memory Model That Permanently Remembers You Without Bolt-On RAG. The core argument is radical — stop bolting on RAG, and make memory a native capability of the model.
The video presented a three-axis comparison:
| Dimension | External Memory (RAG-style) | Native Memory (Metis) | |---|---|---| | Architecture | External (decoupled) | Native (coupled) | | Optimization | Blocked gradients | End-to-end gradients | | Efficiency | Sequential | Parallel |
In plain terms: RAG is a three-stage "retrieve → stuff into prompt → generate" pipeline where the memory module and the main model are decoupled, gradients are cut off, and execution is sequential. Metis instead welds memory into the main model at the parameter level, trained end-to-end with parallel decoding.
---
A few points I personally care about:
1. Is "native" just another form of RAG? Early MemGPT also claimed native memory, but in essence it was hierarchical context window management. If Metis truly achieves "parameter-level coupling," then memory is no longer text but weights — which means every user's memory must be fine-tuned into the parameters. How do you handle the cost? How does personalization work?
2. "End-to-end gradients" sounds great, but how is long-term credit assignment solved? How much does today's conversation contribute to an answer three months from now? That's an extremely long causal chain. RAG sidesteps this problem with explicit retrieval; the native approach currently shows no clean solution.
3. An interesting data point The video had 691 views / 57 favorites — an unusually high 8.2% favorite rate (typical Bilibili videos from creators with ~1k followers see 1–3%). The tech community is clearly watching this direction, but going from paper to product-grade deployment (millions of DAU) still requires two or three orders of magnitude of engineering work.
---
The paper's author put the original on Quark drive: https://pan.quark.cn/s/38c973a1bfef (I couldn't download it — anyone who has, please drop a summary in the comments).
Do you think the native memory path can actually work? Or will RAG's "bolt-on philosophy" survive the way microservices did — theoretically inelegant, but it runs in practice?