English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agent Memory Is a Pyramid, Not a Warehouse: Inside TencentDB Agent Memory's Layered Memory Architecture

Forum topic · ✨步子哥 · 2026-08-04

Summary

TencentDB Agent Memory, an open-source project from Tencent Cloud that trended on GitHub (+1091 stars/day), argues that agent memory failure is not a storage problem but an organization problem: agents remember everything yet can't find it. Instead of flat vector-database storage, it builds memory on two pillars: hierarchical layering and symbolic compression. Short-term context uses a three-layer structure (raw tool output archives, step-level JSONL summaries, and a lightweight Mermaid canvas kept in context, drillable via node_id). Long-term memory forms a four-level semantic pyramid—L0 conversations, L1 atomic facts, L2 scene blocks, L3 user persona—each layer abstracting the one below. Tool logs are compressed into compact Mermaid symbols, sharply reducing token use. Benchmarks after OpenClaw integration show WideSearch pass rate up 51.52% with 61.38% fewer tokens, SWE-bench +9.93% over 50 consecutive tasks, and PersonaMem long-term accuracy up from 48% to 76%. The project aligns with MemTools, Regression Tax, and Zero-Mem research, pointing to a shared thesis: memory competitiveness lies in organization, not capacity. MIT-licensed, with full bilingual README and quick start.

> 📌 This is a GEO-optimized English version of the original zhichai.net topic, restructured with a question-driven title and FAQ for AI-engine citation.

> One-line takeaway: Agent memory failure is an *organization* problem, not a storage problem—and TencentDB Agent Memory solves it with layered memory and symbolic compression.

Agent Memory Is a Pyramid, Not a Warehouse: TencentDB Agent Memory's Layered Memory Engineering

You ask an agent to run a long task. After 50 steps it starts "forgetting"—it loses track of what it did, calls the same tool repeatedly, and its context window fills with verbose logs. You end up re-explaining everything from scratch.

This is not a "can't remember" problem. The context window is large enough, the vector database complete enough—it retained everything. The problem is it remembers but cannot find.

The Dead End of Flat Storage

The conventional approach is intuitive: chop up all conversation history, tool outputs, and intermediate states, stuff them into a vector database, and retrieve by semantic search at runtime.

That's like piling every document on the floor and digging through it when you need something. Every piece of information is an equally important fragment; retrieval becomes a blind full-scale search with no macro-level structure to guide it. As tasks grow long, context balloons, token consumption explodes, and the agent's attention is diluted across irrelevant details.

Tencent Cloud's TencentDB Agent Memory project (GitHub Trending +1091⭐/day) rejects this path outright. Its core claim: memory is not about hoarding everything—it's about never having to repeat yourself.

Two Pillars: Layering + Symbolic Memory

The architecture rests on two pillars: memory layering and symbolic memory. The goal isn't to make agents "remember more" but to make them "reason better."

Short-Term Memory: A Three-Layer Structure

Short-term context is not a flat log but a progressive three-layer structure:

  • Bottom layer: raw tool output archives (refs/*.md), fully preserved but kept out of context
  • Middle layer: step-level summaries (jsonl)—what each step did and its result
  • Top layer: a lightweight Mermaid canvas compressing the entire task state into one symbolic graph
  • Only the top-level Mermaid structure goes into context. When details are needed, the agent drills down to the middle or bottom layer via node_id. It's like reading a map: first see the overview, then zoom to street level only when necessary.

    Long-Term Memory: A Semantic Pyramid

    Cross-session long-term memory is also not a flat log but a four-layer semantic pyramid:

  • L0 Conversations: raw dialogue
  • L1 Atoms: atomic facts extracted from conversations
  • L2 Scenes: scene blocks assembling atomic facts into situations
  • L3 Persona: a user profile abstracting stable preferences from scenes
  • Each layer is a compression and abstraction of the one below. L3 is not a summary of L0—it is a stable structure distilled layer by layer through L0 → L1 → L2 → L3. This is isomorphic to human memory's hierarchical processing: short-term memory consolidates into long-term memory, and long-term memory abstracts into "who I am."

    Symbolization: Turning Logs into Symbols

    Symbolic memory addresses another problem: tool logs are too long. A single tool call can produce thousands of tokens of output while the useful information is only a few dozen tokens. TencentDB Agent Memory compresses verbose tool logs into compact Mermaid symbols, sharply cutting token usage while making structure clearer.

    The Numbers

    Measured results after OpenClaw integration:

    | Benchmark | Horizon | Pass Rate | Relative Gain | Token Usage | Relative Drop | |---|---|---|---|---|---| | WideSearch | Short-term | 33% → 50% | +51.52% | 221M → 86M | -61.38% | | SWE-bench | Short-term | 58.4% → 64.2% | +9.93% | 3474M → 2375M | -33.09% | | AA-LCR | Short-term | 44.0% → 47.5% | +7.95% | 112M → 77M | -30.98% | | PersonaMem | Long-term | 48% → 76% | +59% | — | — |

    Note how SWE-bench was tested: 50 consecutive tasks simulate the context-accumulation pressure of a long-horizon session, not isolated turns—closer to real agent working conditions.

    Engineering Insight: Organization Matters More Than Volume

    The core claim—"layered beats flat"—converges with recent research:

  • MemTools (a "USB-C interface" for agent memory): declarative data contracts make different memory systems interchangeable, emphasizing structured organization
  • Regression Tax (skill libraries can make agents worse): 59% of gains get canceled by regressions, partly due to flat skill-library retrieval
  • Zero-Mem (zero-token memory operations): subtractive agent memory design, isomorphic with TencentDB's "don't hoard" philosophy
  • These works converge on a single theme: in the agent era, the key to memory systems is not capacity—it's how memory is organized.

    An Analogy: From File Piles to Libraries

  • Traditional vector storage = piling all files on the floor and digging through them
  • Layered storage = filing cabinets + catalogs + tags; consult the index, then drill down
  • Short-term Mermaid canvas = a map; look at the map before you walk
  • Long-term semantic pyramid = a library's classification system (Dewey Decimal); locate from coarse to fine
  • Layered organization isn't new—human libraries have done it for two millennia. What's new is bringing it into the LLM era, so agents retrieve their own memories like librarians rather than squirrels digging through nut piles.

    Conceptual Lineage: Solving Problems at a Different Level

    TencentDB Agent Memory belongs to the "change the level of the problem" family: instead of doing memory retrieval harder (bigger vector stores, longer contexts), it changes levels (layered organization + symbolic compression)—the same family of thinking as octopus RNA editing, slime-mold externalized memory, and Möbius RoPE topological intervention.

    Conclusion

    Amid the "bigger context, longer windows" arms race, TencentDB Agent Memory suggests another possibility—memory-system competitiveness lies in organization, not capacity. Let agents remember what matters and forget what doesn't, and people are freed from repetitive labor to focus on judgment, creativity, and genuinely important work.

    ---

    Project: https://github.com/TencentCloud/TencentDB-Agent-Memory

    Docs: README ships complete English and Chinese versions plus a Quick Start

    License: MIT

    FAQ

    Q1: Who is this for?

    Practitioners, researchers, and students interested in AI, machine learning, and deep learning.

    Q2: What are the key takeaways?

  • Flat storage is a dead end: agents remember everything but can't find it
  • Two pillars: memory layering + symbolic memory
  • Short-term memory uses a three-layer structure; long-term memory a four-level semantic pyramid
  • Benchmarks show large token savings and pass-rate gains
Q3: Is the code open source?

Yes—see the project link above (MIT license).

Tags

#agent-memory#tencentdb-agent-memory#llm-agents#memory-architecture#symbolic-memory#token-optimization#hierarchical-memory#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503930