Agent Memory Is a Pyramid, Not a Warehouse: Hierarchical Memory Engineering in TencentDB Agent Memory
Run an agent on a long task, and after ~50 steps it starts "forgetting"—it forgets what it did, repeats the same tool calls, and the context window fills with verbose logs. You have to re-explain everything from scratch.
This is not a problem of not remembering. With large context windows and complete vector databases, it remembers everything. The problem is: it remembers but cannot find.
The Deadlock of Flat Storage
The conventional approach is intuitive: chop up all conversation history, tool outputs, and intermediate states, stuff them into a vector database, and do semantic retrieval when needed.
This is like piling all files on the floor and crouching down to dig through them. Every piece of information is an equally weighted fragment; retrieval becomes blind, exhaustive searching with no macro structure to guide it. As tasks grow long, context bloats, token consumption explodes, and the agent's attention is diluted across irrelevant details.
Tencent Cloud's TencentDB Agent Memory project (GitHub Trending +1091⭐/day) rejects this path outright. Its core claim: memory is not about hoarding everything; it is about never having to repeat yourself.
Two Pillars: Layering + Symbolization
The architecture rests on two pillars: memory layering and symbolic memory. The goal is not to make the agent "remember more" but to make it "reason better."
Three-Tier Short-Term Memory
Short-term context is not a flat log but a three-tier progressive structure:
- Bottom tier: archived raw tool outputs (
refs/*.md)—kept in full, but never enters context - Middle tier: step-level summaries (
jsonl)—what each step did and its result - Top tier: a lightweight Mermaid canvas compressing the entire task state into one symbolic graph
- L0 Conversation: raw dialogue
- L1 Atomic: atomic facts extracted from conversations
- L2 Scene: scene blocks assembling atomic facts into situations
- L3 Persona: user profile abstracting stable preferences from scenes
- MemTools (a USB-C interface for agent memory): declarative data contracts make memory systems interchangeable, emphasizing structured organization
- Regression Tax (skill libraries can make agents worse): 59% of gains were canceled by regressions, partly due to flat skill-library retrieval
- Zero-Mem (zero-token memory operations): subtracting from agent memory systems, isomorphic to the "don't hoard" principle
- Traditional vector storage = all files piled on the floor, dug through on demand
- Layered storage = filing cabinets + catalog + tags; consult the catalog, then drill down
- Short-term Mermaid canvas = a map; check the map before walking
- Long-term semantic pyramid = a library classification system (Dewey Decimal), locating from coarse to fine
Only the top-level Mermaid structure enters the context. When details are needed, the agent drills down via node_id to the middle or bottom tier. It is like reading a map: first get your bearings, then zoom into street view when needed.
The Semantic Pyramid of Long-Term Memory
Cross-session long-term memory is also not a flat log but a four-level semantic pyramid:
Each level is a compression and abstraction of the one below. L3 is not a summary of L0—it is distilled layer by layer from L0 → L1 → L2 → L3 into a stable structure. This is isomorphic to hierarchical human memory processing: short-term memory consolidates into long-term memory; long-term memory abstracts into "who I am."
Symbolization: Turning Logs into Symbols
Symbolic memory solves another problem: tool logs are too long. A single tool call can emit thousands of tokens while carrying only tens of tokens of useful information. TencentDB Agent Memory compresses verbose tool logs into compact Mermaid symbols, drastically cutting token usage while improving structure.
The Data Speaks
Measured results after integration with OpenClaw:
| Benchmark | Task | Pass Rate | Relative Gain | Token Usage | Relative Drop | |---|---|---|---|---|---| | WideSearch | short-term | 33% → 50% | +51.52% | 221M → 86M | -61.38% | | SWE-bench | short-term | 58.4% → 64.2% | +9.93% | 3474M → 2375M | -33.09% | | AA-LCR | short-term | 44.0% → 47.5% | +7.95% | 112M → 77M | -30.98% | | PersonaMem | long-term | 48% → 76% | +59% | — | — |
Note how SWE-bench was tested: 50 consecutive tasks simulate the context-accumulation pressure of long-horizon sessions, not isolated turns—closer to real agent working conditions.
Engineering Insight: Organization Beats Volume
The core claim—"layered beats flat"—converges with recent research:
These works point to a common theme: in the agent era, the key to memory systems is not capacity but organization.
Analogy: From File Piles to Libraries
Conceptual Genealogy: Solving at a Different Level
TencentDB Agent Memory belongs to the "change the level of the problem" family—not stronger memory retrieval (bigger vector stores, longer contexts), but a different level entirely (hierarchical organization + symbolic compression). Same family of thinking as octopus RNA editing, slime-mold externalized memory, and Möbius RoPE topological intervention.
Conclusion
Amid the arms race of "bigger context, longer windows," TencentDB Agent Memory points to another possibility: the competitiveness of a memory system lies not in capacity but in organization. Let the agent remember what matters and forget what doesn't, and people are freed from repetitive work to focus on judgment, creativity, and what really counts.
---
Project: https://github.com/TencentCloud/TencentDB-Agent-Memory Docs: README includes full English and Chinese versions plus a Quick Start License: MIT