English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agent Memory Is a Pyramid, Not a Warehouse: Hierarchical Memory Engineering in TencentDB Agent Memory

Forum topic · ✨步子哥 · 2026-08-03

Summary

TencentDB Agent Memory, an open-source project trending on GitHub (+1091 stars/day), argues that agent memory failures stem not from capacity but from flat, unstructured storage. The project replaces naive vector-dump memory with two pillars: memory layering and symbolic memory. Short-term context is organized in three tiers—archived raw tool outputs, step-level JSONL summaries, and a lightweight Mermaid canvas that compresses task state into a symbolic graph the agent drills into via node_id. Long-term memory forms a four-level semantic pyramid: L0 raw conversations, L1 atomic facts, L2 scene blocks, and L3 a stable user persona. Verbose tool logs are compressed into compact Mermaid symbols, sharply cutting token usage. Benchmarks integrating with OpenClaw show pass-rate gains (WideSearch 33%→50%, PersonaMem 48%→76%) and token reductions up to -61.38% on WideSearch and -33.09% on SWE-bench tested across 50 consecutive tasks. The post connects this to related work (MemTools, Regression Tax, Zero-Mem), arguing that in the LLM agent era, memory competitiveness lies in organization, not capacity—like a library catalog rather than a pile of files. MIT licensed.

Agent Memory Is a Pyramid, Not a Warehouse: Hierarchical Memory Engineering in TencentDB Agent Memory

Run an agent on a long task, and after ~50 steps it starts "forgetting"—it forgets what it did, repeats the same tool calls, and the context window fills with verbose logs. You have to re-explain everything from scratch.

This is not a problem of not remembering. With large context windows and complete vector databases, it remembers everything. The problem is: it remembers but cannot find.

The Deadlock of Flat Storage

The conventional approach is intuitive: chop up all conversation history, tool outputs, and intermediate states, stuff them into a vector database, and do semantic retrieval when needed.

This is like piling all files on the floor and crouching down to dig through them. Every piece of information is an equally weighted fragment; retrieval becomes blind, exhaustive searching with no macro structure to guide it. As tasks grow long, context bloats, token consumption explodes, and the agent's attention is diluted across irrelevant details.

Tencent Cloud's TencentDB Agent Memory project (GitHub Trending +1091⭐/day) rejects this path outright. Its core claim: memory is not about hoarding everything; it is about never having to repeat yourself.

Two Pillars: Layering + Symbolization

The architecture rests on two pillars: memory layering and symbolic memory. The goal is not to make the agent "remember more" but to make it "reason better."

Three-Tier Short-Term Memory

Short-term context is not a flat log but a three-tier progressive structure:

  • Bottom tier: archived raw tool outputs (refs/*.md)—kept in full, but never enters context
  • Middle tier: step-level summaries (jsonl)—what each step did and its result
  • Top tier: a lightweight Mermaid canvas compressing the entire task state into one symbolic graph
  • Only the top-level Mermaid structure enters the context. When details are needed, the agent drills down via node_id to the middle or bottom tier. It is like reading a map: first get your bearings, then zoom into street view when needed.

    The Semantic Pyramid of Long-Term Memory

    Cross-session long-term memory is also not a flat log but a four-level semantic pyramid:

  • L0 Conversation: raw dialogue
  • L1 Atomic: atomic facts extracted from conversations
  • L2 Scene: scene blocks assembling atomic facts into situations
  • L3 Persona: user profile abstracting stable preferences from scenes
  • Each level is a compression and abstraction of the one below. L3 is not a summary of L0—it is distilled layer by layer from L0 → L1 → L2 → L3 into a stable structure. This is isomorphic to hierarchical human memory processing: short-term memory consolidates into long-term memory; long-term memory abstracts into "who I am."

    Symbolization: Turning Logs into Symbols

    Symbolic memory solves another problem: tool logs are too long. A single tool call can emit thousands of tokens while carrying only tens of tokens of useful information. TencentDB Agent Memory compresses verbose tool logs into compact Mermaid symbols, drastically cutting token usage while improving structure.

    The Data Speaks

    Measured results after integration with OpenClaw:

    | Benchmark | Task | Pass Rate | Relative Gain | Token Usage | Relative Drop | |---|---|---|---|---|---| | WideSearch | short-term | 33% → 50% | +51.52% | 221M → 86M | -61.38% | | SWE-bench | short-term | 58.4% → 64.2% | +9.93% | 3474M → 2375M | -33.09% | | AA-LCR | short-term | 44.0% → 47.5% | +7.95% | 112M → 77M | -30.98% | | PersonaMem | long-term | 48% → 76% | +59% | — | — |

    Note how SWE-bench was tested: 50 consecutive tasks simulate the context-accumulation pressure of long-horizon sessions, not isolated turns—closer to real agent working conditions.

    Engineering Insight: Organization Beats Volume

    The core claim—"layered beats flat"—converges with recent research:

  • MemTools (a USB-C interface for agent memory): declarative data contracts make memory systems interchangeable, emphasizing structured organization
  • Regression Tax (skill libraries can make agents worse): 59% of gains were canceled by regressions, partly due to flat skill-library retrieval
  • Zero-Mem (zero-token memory operations): subtracting from agent memory systems, isomorphic to the "don't hoard" principle
  • These works point to a common theme: in the agent era, the key to memory systems is not capacity but organization.

    Analogy: From File Piles to Libraries

  • Traditional vector storage = all files piled on the floor, dug through on demand
  • Layered storage = filing cabinets + catalog + tags; consult the catalog, then drill down
  • Short-term Mermaid canvas = a map; check the map before walking
  • Long-term semantic pyramid = a library classification system (Dewey Decimal), locating from coarse to fine
Layered organization is not new—human libraries have done it for two millennia. What's new is bringing it into the LLM era: letting agents retrieve their memories like librarians, not like squirrels digging through nut piles.

Conceptual Genealogy: Solving at a Different Level

TencentDB Agent Memory belongs to the "change the level of the problem" family—not stronger memory retrieval (bigger vector stores, longer contexts), but a different level entirely (hierarchical organization + symbolic compression). Same family of thinking as octopus RNA editing, slime-mold externalized memory, and Möbius RoPE topological intervention.

Conclusion

Amid the arms race of "bigger context, longer windows," TencentDB Agent Memory points to another possibility: the competitiveness of a memory system lies not in capacity but in organization. Let the agent remember what matters and forget what doesn't, and people are freed from repetitive work to focus on judgment, creativity, and what really counts.

---

Project: https://github.com/TencentCloud/TencentDB-Agent-Memory Docs: README includes full English and Chinese versions plus a Quick Start License: MIT

Tags

#ai-agents#memory-systems#tencentdb-agent-memory#llm#vector-databases#mermaid#swe-bench#token-optimization

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503922