English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Context Engineering for AI Agents: Sessions, Memory, and the Context Pipeline Explained

Forum topic · ✨步子哥 · 2025-12-28

Summary

This in-depth guide treats agent context as a production pipeline rather than a static prompt. Because LLMs are stateless, durable agent behavior requires two externalized systems: Sessions (per-user, per-conversation event logs plus working state) and Memory (long-term, distilled facts and skills that persist across sessions). The article walks through the four-step per-turn loop—Fetch, Prepare, Invoke, Upload—stressing that prompt assembly is a blocking hot path while memory writes should be asynchronous. It covers long-conversation compression (sliding window, token-based truncation, recursive summarization, and trigger strategies), memory generation as an LLM-driven ETL (extraction, consolidation, provenance tracking, forgetting), and retrieval scoring that combines relevance, recency, and importance. Further topics include multi-agent session patterns (shared unified history vs. separate individual histories), why a framework-agnostic Memory layer enables cross-framework interoperability, declarative vs. procedural knowledge, injection points for memories, three-tier evaluation (generation quality, retrieval metrics, end-to-end task success), and privacy requirements such as strict user isolation, PII masking before storage, and defense against memory poisoning.

Treating "context" as an assembly line: how Sessions and Memory let agents remember, run fast, and stay within bounds.

If tools give an Agent its "hands," then Context Engineering is the methodology that keeps those hands from grabbing blindly — it decides what the model sees, ignores, remembers, and forgets each turn. Models are naturally stateless: once a call ends, it wakes up with no memory of what just happened. To give an Agent continuous conversation, long-term personalization, and cross-session experience accumulation, you must externalize "state" into two systems: Session (the conversation workbench) and Memory (the long-term filing cabinet), dynamically assembling them into the context window each turn.

This article focuses on engineering practice: how to design Sessions, how to compress long conversations, how Memory is generated/consolidated/retrieved, and why Memory becomes the "common layer" for multi-agent and cross-framework collaboration. Finally, privacy/security and evaluation metrics are brought in — because "remembering" is not the goal; remembering correctly, retrieving reliably, using stably, and not leaking is.

Chapter 1: What Is Context Engineering — From "Writing Prompts" to "Assembling the Full Request"

Traditional Prompt Engineering writes a fixed system instruction; Context Engineering concerns the complete payload of every call: dynamically constructing a "stateful" request from the user, session, tool results, external knowledge, and long-term memory.

Think of it as *mise en place* before cooking: not just a recipe, but fresh ingredients, knives, seasoning, and plating requirements prepared in advance — the model doesn't need to "guess," it reasons and acts with maximum certainty in minimal noise.

Context typically consists of three categories:

A. Context that guides reasoning (behavioral constitution)

  • System Instructions: persona, capability boundaries, rules
  • Tool Definitions: tool schemas and descriptions
  • Few-shot Examples: examples guiding reasoning paths
  • B. Evidence and factual data (the evidence chain)

  • Long-Term Memory: user info/experience accumulated across sessions
  • External Knowledge: documents/database info retrieved via RAG
  • Tool Outputs
  • Sub-Agent Outputs
  • Artifacts: files, images, and other non-text materials
  • C. Current interaction (the matter at hand)

  • Conversation History
  • State / Scratchpad: temporary variables, shopping cart, draft calculations
  • User Prompt
  • > Tip: more context is not better. > Long context brings cost and latency, plus "context rot": the more information, the easier the model loses the key points and quality degrades. Context Engineering aims for "just enough, not too much."

    Chapter 2: The Per-Turn "Context Pipeline" — Fetch → Prepare → Invoke → Upload

    A production Agent typically loops through four steps each turn:

    1. Fetch Context: retrieve needed context (memory, RAG, recent events) 2. Prepare Context: build the final prompt (this is the blocking hot path) 3. Invoke LLM and Tools: model and tool iterations append outputs to context 4. Upload Context: write new information to persistent storage (usually async in the background)

    The most overlooked but critical point: Prepare is a blocking hot path; Upload should be backgrounded. Putting expensive distillation/consolidation work into the hot path makes the experience "lag on every message."

    Chapter 3: Sessions — The Conversation Workbench (Event Log + Working State)

    3.1 A Session consists of Events and State

  • Events: time-ordered conversation events: user input, agent replies, tool calls, tool outputs
  • State: structured working memory (scratchpad): cart contents, task progress, confirmed parameters
  • A Session is a single-user, single-conversation container. Users can have multiple sessions, but they are disconnected "project workbenches" that don't naturally share memory (sharing requires Memory).

    3.2 Session storage in production: runtimes are usually stateless

    Most Agent runtimes are stateless: memory is cleared when a request ends. So session history must be written to persistent storage (database / managed session service) and pulled back at the start of each turn.

    Chapter 4: Framework Differences and Multi-Agent Session Patterns — Shared Big Ledger vs. Separate Small Ledgers

    4.1 The essence of framework differences: unified internals, diverse external protocols

    Frameworks are essentially "translators": developers use the framework's internal event structures, and the framework maps them to what each model API requires. This decouples model vendors but creates cross-framework "semantic isolation."

    4.2 Two patterns for organizing session history in multi-agent systems

    The key is "how information is shared."

    #### A. Shared Unified History All agents write messages, tool calls, and observations into the same time-series log. Suitable for tightly coupled collaboration and pipeline tasks depending on a "single source of truth."

    Pros: globally traceable, natural handoffs; Cons: history bloat, noise, harder permission isolation.

    #### B. Separate Individual Histories Each agent has private history and only exposes final results (like a black-box tool), often via "Agent-as-a-Tool" or A2A messages exchanging results, not process.

    Pros: clear boundaries, low leakage risk; Cons: thin shared context, collaboration needs extra design.

    Chapter 5: The Hard Problem of Cross-Framework Interop — Sessions Aren't Portable, Memory Can Be the "Common Layer"

    Different frameworks' session/event storage models are tightly bound to internal object structures, so a Session written by one framework is hard for another to read directly. A2A messages work, but "rich state" still needs a translation layer.

    The more robust pattern: abstract shared knowledge into a framework-agnostic Memory layer. Memory stores not framework event objects but distilled facts/entities/summaries (strings or dicts), making it a shared cognitive resource across frameworks and agents.

    Chapter 6: Production Considerations for Sessions — Isolation, Integrity, Performance, and Compression

    6.1 Security and privacy: strict isolation + PII masking before storage

  • Sessions belong to one user; ACLs must strictly prevent cross-user access
  • Best practice: mask/remove PII before writing to storage, reducing blast radius and aiding GDPR/CCPA compliance
  • 6.2 Data integrity: ordering consistency + lifecycle (TTL)

  • Event append order must be deterministic (log disorder breaks reasoning)
  • Sessions shouldn't be kept forever: use TTL or archival policies to control cost and compliance risk
  • 6.3 Performance: sessions sit on the hot path — must be "fast and small"

    Each turn pulls session history to build the prompt; larger history means slower, costlier turns. Key optimization: filter/compress history before sending to the model (e.g., strip stale tool outputs).

    Chapter 7: Long Conversation Management — Packing History Like Luggage

    Four hard constraints: 1) Context window limits 2) Token cost 3) Latency 4) Quality (noise and autoregressive error)

    Compression strategies (simple to complex):

    7.1 Sliding Window: keep the last N turns

    Simple and effective, but may lose early key constraints.

    7.2 Token-Based Truncation: truncate by token budget

    Better fits cost control; same risk of cutting key facts.

    7.3 Recursive Summarization: old content becomes summaries

    Replace older segments with summaries coexisting with recent raw turns. Engineering points:
  • Generate summaries asynchronously in the background and persist them; avoid re-summarizing every turn
  • Track which events are already covered by summaries to avoid re-injecting raw text
  • 7.4 Compression triggers: when?

  • Count-based: compress when turns/tokens exceed a threshold (most common, "good enough")
  • Time-based: compress in the background after user inactivity
  • Event-based: compress when a subtask/topic ends
  • Chapter 8: Memory — The Long-Term Filing Cabinet (Distilled, Reusable, "Less but Better")

    Memory is not "storing the conversation" but extracting and solidifying valuable information that persists across sessions, used for personalization, context management, data insights, and self-improvement.

    Four capabilities Memory provides

  • Personalization (preferences, facts, history)
  • Replacing long histories (summaries/key facts reduce tokens)
  • Insight at scale (aggregate trend mining, with privacy protection)
  • Self-improvement (recording successful strategies/tool paths as playbooks)
  • Chapter 9: Memory vs. RAG — One Knows the World, the Other Knows the User

  • RAG: injects external, static, authoritative facts; generally shared, read-only
  • Memory: accumulates user-specific, dynamic, isolated context; needs writing and evolution
  • One sentence: RAG makes the system a "facts expert"; Memory makes it a "user expert."

    Chapter 10: Memory Structure and Types — Content + Metadata; "Knowing What" and "Knowing How"

    10.1 Basic structure

  • content: fact fragments (text or structured JSON/dict)
  • metadata: id, owner, tags, source, etc.
  • 10.2 Knowledge types: Declarative vs. Procedural

  • Declarative (knowing what): facts/preferences/events
  • Procedural (knowing how): skills/processes/workflow "playbooks"
  • > Key difference: procedural memory is not "retrieving data" but "retrieving an executable plan" — more like a reasoning-augmentation layer. Compared with fine-tuning (offline, weight changes), procedural memory injects playbooks online, correcting behavior quickly via in-context learning.

    Chapter 11: Organizing and Storing Memory — Collections / User Profiles / Rolling Summaries; Vector DBs / Knowledge Graphs / Hybrid

    11.1 Organization patterns

  • Collections: piles of atomic memories, good for search and multi-topic
  • Structured User Profile: stable fields (name, preferences), fast reads
  • Rolling Summary: one continuously updated summary, often used to compress long sessions
  • 11.2 Storage architectures

  • Vector databases: semantic similarity retrieval for unstructured memories
  • Knowledge graphs: entity-relation reasoning for structured associations
  • Hybrid: graph nodes with embeddings, supporting both semantic and relational retrieval
  • Chapter 12: Memory Generation — LLM-Driven ETL of Extraction + Consolidation

    Memory generation isn't simple summarization but an LLM-driven ETL:

    1) Ingestion: input source data (usually session history) 2) Extraction & Filtering: extract meaningful information per "topic definitions" (no match, no memory) 3) Consolidation: dedup, conflict resolution, update/create/delete, and forgetting (TTL/low confidence) 4) Storage: write to vector DBs/graphs and other persistence layers

    12.1 Extraction: deciding "what's worth remembering"

    "Meaningful" depends entirely on business goals. Common implementations:
  • Schema/template extraction (structured output)
  • Natural language topic definitions
  • Few-shot examples (especially effective for subtle domains)
  • Many systems use "rolling summaries" as an auxiliary extraction input, improving efficiency and avoiding rescanning full history each turn.

    12.2 Consolidation: without it, memory becomes a contradictory garbage heap

    Consolidation must handle:
  • Duplicates (same fact, multiple expressions)
  • Conflicts (changing user preferences)
  • Evolution (facts becoming more specific)
  • Decay and forgetting (old, low-confidence items get deleted/downweighted)
  • Common operations: UPDATE / CREATE / DELETE (or INVALIDATE).

    Chapter 13: Memory Provenance — Why Should You Trust It?

    Long-term memory without source tracking drifts toward "confident nonsense." Provenance should at least record:

  • Source type (preloaded system data / user input / tool output)
  • Freshness (time)
  • Confidence changes (rising with multi-source corroboration, decaying over time)
  • An important recommendation: generating long-term memory from tool outputs is generally not recommended — tool data suits short-term caching and is easily stale and fragile.

    Engineering value of provenance:

  • Trust tiers during consolidation (trusted sources first, recent first, multi-source corroboration first)
  • When deleting a data source, precisely regenerate via lineage instead of wiping everything
  • Chapter 14: When to Generate Memory — Memory-as-a-Tool Is Smarter but Costs Must Be Controlled

    Trigger strategy trade-off is "cost vs. fidelity":

  • Session end: cheap, but possibly low fidelity
  • Every N turns: common middle ground
  • Every turn, real-time: freshest, most expensive
  • Explicit user instruction: controllable but incomplete coverage
  • Memory-as-a-Tool: the Agent itself decides when to call create_memory — smarter, but needs stricter tool definitions and policy controls to avoid "over-remembering"
  • Background vs. blocking: memory generation should almost always be async

    Memory generation requires LLM calls and DB writes and must be stripped from the hot path: the user gets a reply first, extraction/consolidation happens in the background. Otherwise the experience becomes unusable.

    Chapter 15: Memory Retrieval — Relevance Isn't Enough; Recency and Importance Matter Too

    The goal: find the most useful memories within a strict latency budget.

    Common scoring dimensions:

  • Relevance (semantic)
  • Recency (temporal freshness)
  • Importance (annotated at generation time)
  • Relying only on vector similarity is a common pitfall: it surfaces memories that are "similar but old/trivial." A multi-dimensional hybrid score is more robust.

    Advanced enhancements (but slower):

  • Query rewriting (one extra LLM call)
  • Reranking (take top 50, rerank with an LLM)
  • Training a dedicated retriever (needs labeled data, costly)
  • Pair with caching to avoid repeating expensive pipelines.

    Retrieval timing: proactive prefetch vs. reactive query

  • Proactive prefetch: fetch at the start of every turn; simple but adds latency; cacheable
  • Reactive (Memory-as-a-Tool): query only when needed; saves cost but may add an LLM call; the Agent must "know what memory types might exist," otherwise it won't know when to query
  • Chapter 16: Putting Memory into Context — Where You Put It Is Where the Weight Lands

    Three main injection methods:

    16.1 Into system instructions

    Pros: high authority, clean conversation history; good for stable info (user profiles). Risks: over-influence — the model may force every topic toward the memory; system instructions are usually text-only (bad for multimodal memories); incompatible with "let the model decide whether to call a memory tool."

    16.2 Injected into conversation history

    Pros: flexible; good for situational memories. Risks: noisy, high token cost, possible "conversation injection" illusion (the model treats memory as something the user just said). If injecting user-level memories with the user role, keep first-person perspective consistency.

    16.3 As tool output

    With reactive querying, memories enter context as tool output, naturally landing in the conversation sequence. Pros: clear provenance; Cons: poor retrieval quality directly pollutes reasoning.

    Chapter 17: Evaluating Whether the Memory System Actually Works — Quality, Retrieval, End-to-End

    Evaluation has three layers:

    17.1 Generation quality: is it remembered correctly?

    Compare against a human golden set:
  • Precision: how much recorded is correct and relevant (guards against "over-remembering" pollution)
  • Recall: are key facts missed (guards against "critical fact loss")
  • F1: combined metric
  • 17.2 Retrieval performance: can it be found, and fast?

  • Recall@K: does the needed memory appear in topK
  • Latency: retrieval within a strict budget (e.g., <200ms)
  • (Complex rewriting/reranking generally doesn't fit the real-time hot path unless cacheable and not easily stale.)

    17.3 End-to-end task success: does memory improve task completion?

    Use an LLM judge to compare final answers with golden answers, determining whether memory truly improves the target outcome.

    Chapter 18: Privacy and Security — Memory Is a "Corporate Archive," Not a "Casual Notepad"

    The bottom line for long-term memory:

  • Strict isolation (user/tenant scope): cross-user leakage is a fatal incident
  • User control: opt-out and deletion supported
  • PII masked before storage
  • Defense against memory poisoning: validate and sanitize injected/forged information
  • Cross-user shared procedural memory must be strongly anonymized, or "experience reuse" becomes "information leakage"

References (5 items)

1. Retrieval-Augmented Generation overview: https://cloud.google.com/use-cases/retrieval-augmented-generation?hl=en 2. Agent Engine Sessions overview: https://cloud.google.com/vertex-ai/generative-ai/docs/agent-engine/sessions/overview 3. Agent Engine Memory Bank – generate memories: https://cloud.google.com/vertex-ai/generative-ai/docs/agent-engine/memory-bank/generate-memories 4. Model Armor overview: https://cloud.google.com/security-command-center/docs/model-armor-overview 5. Gemini long context limitations: https://ai.google.dev/gemini-api/docs/long-context#long-context-limitations

Tags

#context-engineering#ai-agents#session-management#agent-memory#rag#llm#multi-agent-systems#privacy

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/176415196