English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Zero-Mem: Zero-Token Memory Operations for AI Agents by Removing Generation

Forum topic · ✨步子哥 · 2026-08-04

Summary

Zero-Mem is a memory system for AI agents that performs all memory operations—summarization, extraction, updating, and retrieval—without a single LLM call. Instead of using LLMs to generate memory abstractions, it preserves raw interaction traces and builds two lightweight views: an entity-context graph with Personalized PageRank retrieval for relational queries, and a turn/window/episode/local time hierarchy for temporal queries. A lightweight query analyzer (not an LLM) routes questions to the appropriate view, and deterministic calibration steps filter and verify evidence. Benchmarks report zero LLM tokens for memory operations, 57.6% lower memory-operation latency versus the fastest baseline (with the same QA reader and context budget), and competitive answer quality. Ablations show both views and calibration are necessary. The system also preserves provenance—original text, source IDs, and timestamps—enabling auditable answers. Paper: arxiv.org/abs/2607.29377; code: github.com/TheMoon0815/Zero-mem.

Zero-Mem: Zero-Token Memory Operations for AI Agents by Removing Generation

Ever asked an AI assistant what coffee flavor you mentioned three days ago, only to watch it scan chat history—an action that itself burns a large amount of tokens? Every time it "recalls," it invokes an LLM to summarize, extract, and retrieve past conversations. The more memory accumulates, the more expensive recall becomes.

Zero-Mem asks a radical question: do memory operations really need LLM calls at all?

Paper: https://arxiv.org/abs/2607.29377 Code: https://github.com/TheMoon0815/Zero-mem

The "Generation Tax" of Existing Agent Memory Systems

Mainstream agent memory systems—Zep, Mem0, A-Mem, MemoryOS, GAM—all use LLMs to manage memory:

1. Summarization: compressing long conversations into key points 2. Extraction: pulling entities, relations, and events from dialogue 3. Updating: merging, modifying, deleting old memories when new information arrives 4. Retrieval: using an LLM to judge which memories are relevant

Each operation is an LLM call consuming input and output tokens. A long conversation may trigger dozens of memory operations. This is the "generation tax"—you pay not only for the final answer but for the act of remembering.

The deeper problem: generation is loss. Compressing raw dialogue into summaries loses detail; merging conversations blurs timelines; letting an LLM "understand" memory can introduce hallucination. The chain of evidence back to the original source breaks.

The Core Insight: Filter, Don't Generate

> Memory operations don't need generation—they need structured filtering.

It's like library management. The traditional approach: every batch of new books gets read by a librarian (the LLM), who writes summaries, categorizes, and builds index cards. Zero-Mem's approach: put books on the shelf untouched, build a catalog system (entity graph + time hierarchy), and look things up by the catalog to find the original text.

Zero-Mem keeps all raw interaction traces with no generative abstraction, then builds two complementary views:

View 1: Entity–Context Graph (relational view)

A lightweight NER model (e.g., spaCy—not an LLM) scans each dialogue for entities (people, places, organizations). A graph is built with:

  • Entity-to-context edges: an entity appears in a dialogue segment
  • Context-to-context edges: adjacent dialogue segments
  • The graph records observed facts—not LLM-inferred relations. When you ask "how did the project I discussed with Xiao Wang go?", Zero-Mem:

    1. Extracts the entity "Xiao Wang" 2. Locates the entity node 3. Spreads activation along edges via Personalized PageRank (a graph algorithm, not LLM generation) 4. Follows adjacency edges to recover surrounding context

    View 2: Temporal Hierarchy (local view)

    Graphs are good for relations but not temporal order. Zero-Mem organizes dialogue hierarchically:

  • Turn: one complete Q&A interaction
  • Window: several adjacent turns
  • Episode: several windows forming a complete event
  • Local: the immediate context of candidate evidence
  • "What did we discuss around 3 PM yesterday?" prioritizes this view, narrowing from Episode → Window → Turn.

    Dual-View Coordination

    Different questions need different views:

  • "Xiao Wang's project progress" → relational query → graph view prioritized
  • "What we discussed yesterday afternoon" → temporal query → hierarchy view prioritized
  • "Xiao Wang said the project was delayed—how was it resolved?" → hybrid → both views activated
  • A lightweight query analyzer (not an LLM) classifies the question and weights the two views. Fused results enter an evidence closure step that adds relational links and local context.

    Deterministic Calibration: The Last Line of Defense

    Before evidence reaches the answering LLM, two deterministic steps apply:

    1. Evidence calibration: discard conflicting evidence (keep the latest value for the same fact), filter content that doesn't support the question 2. Answer calibration: after the LLM answers, check evidence support, type match, and format

    Both are deterministic rules—no LLM calls.

    Key Results: Zero Tokens + 57.6% Latency Reduction

    Tested on multiple long-memory and long-context QA benchmarks:

  • LLM calls for memory operations: 0 (baselines use dozens)
  • LLM tokens for memory operations: 0 (encoder compute is billed separately)
  • 57.6% lower memory-operation latency vs. the fastest baseline, using the same QA reader and equal context budget
  • Competitive accuracy (not the strongest, but close to baselines on multiple benchmarks)
  • Ablations: Both Views Are Essential

  • Removing the graph view → relational question accuracy drops significantly
  • Removing the time hierarchy → temporal question accuracy drops
  • Removing dual-view coordination → overall performance drops
  • Removing deterministic calibration → answer trustworthiness drops
  • Zero-Mem isn't faster because it does less—it replaces generative methods with structured ones while maintaining quality.

    Deeper Significance

    Zero-Mem touches a fundamental design question: what should the LLM actually do?

    The current trend is "LLM for everything"—memory management, tool selection, result verification—all LLM calls, stacking cost and latency. Zero-Mem proposes the opposite:

    > The LLM should only do what it's best at—understanding questions and generating answers. Every structured, rule-based, algorithmic operation should be done by specialized tools.

    This echoes Euclid-MCP's "let the LLM be the poet, let Prolog be the accountant." Division of labor beats unification—Zero-Mem takes it to the extreme: zero LLM calls for memory operations.

    Another notable point is provenance: every memory unit retains original text, source identifiers, and timestamps, so any final answer is traceable. You can't explain one LLM's answer with another LLM's summary—two stacked hallucinations are unauditable.

    Limitations and Outlook

    1. NER dependency: graph quality depends on spaCy; Chinese and multilingual NER may be weaker 2. Complex reasoning: some memory operations genuinely need inference (e.g., inferring preference shifts over months); pure retrieval may not suffice 3. Limited benchmarks: QA benchmarks only; real agent scenarios (multi-turn tool calls, multimodal input) remain untested

    Still, Zero-Mem's core contribution is proving feasibility: structured memory operations can require zero LLM generation. That proof opens the door for more complex structured operations built on a zero-generation foundation.

    Conclusion

    Zero-Mem recalls an old engineering principle: the best code is no code. Likewise, the best LLM call is no call. When an operation doesn't need an LLM, you save money and gain determinism, auditability, and speed.

    Sometimes all you need is a good catalog system.

    ---

    Paper: https://arxiv.org/abs/2607.29377 Code: https://github.com/TheMoon0815/Zero-mem

    FAQ

    Q1: Who is this for?

    Practitioners, researchers, and students interested in AI agents, memory systems, and LLM efficiency.

    Q2: What are the key takeaways?

  • Existing agent memory systems pay a "generation tax" of LLM calls
  • Zero-Mem replaces generation with structured filtering: an entity-context graph plus a temporal hierarchy
  • Zero LLM tokens for memory operations, 57.6% lower latency, competitive accuracy
Q3: Is the code open source?

Yes: https://github.com/TheMoon0815/Zero-mem

Tags

#ai-agents#memory-systems#zero-mem#llm-efficiency#knowledge-graph#retrieval#rag#provenance

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503925