Treating "context" as an assembly line: how Sessions and Memory let agents remember, run fast, and stay within bounds.
If tools give an Agent its "hands," then Context Engineering is the methodology that keeps those hands from grabbing blindly — it decides what the model sees, ignores, remembers, and forgets each turn. Models are naturally stateless: once a call ends, it wakes up with no memory of what just happened. To give an Agent continuous conversation, long-term personalization, and cross-session experience accumulation, you must externalize "state" into two systems: Session (the conversation workbench) and Memory (the long-term filing cabinet), dynamically assembling them into the context window each turn.
This article focuses on engineering practice: how to design Sessions, how to compress long conversations, how Memory is generated/consolidated/retrieved, and why Memory becomes the "common layer" for multi-agent and cross-framework collaboration. Finally, privacy/security and evaluation metrics are brought in — because "remembering" is not the goal; remembering correctly, retrieving reliably, using stably, and not leaking is.
Chapter 1: What Is Context Engineering — From "Writing Prompts" to "Assembling the Full Request"
Traditional Prompt Engineering writes a fixed system instruction; Context Engineering concerns the complete payload of every call: dynamically constructing a "stateful" request from the user, session, tool results, external knowledge, and long-term memory.
Think of it as *mise en place* before cooking: not just a recipe, but fresh ingredients, knives, seasoning, and plating requirements prepared in advance — the model doesn't need to "guess," it reasons and acts with maximum certainty in minimal noise.
Context typically consists of three categories:
A. Context that guides reasoning (behavioral constitution)
- System Instructions: persona, capability boundaries, rules
- Tool Definitions: tool schemas and descriptions
- Few-shot Examples: examples guiding reasoning paths
- Long-Term Memory: user info/experience accumulated across sessions
- External Knowledge: documents/database info retrieved via RAG
- Tool Outputs
- Sub-Agent Outputs
- Artifacts: files, images, and other non-text materials
- Conversation History
- State / Scratchpad: temporary variables, shopping cart, draft calculations
- User Prompt
- Events: time-ordered conversation events: user input, agent replies, tool calls, tool outputs
- State: structured working memory (scratchpad): cart contents, task progress, confirmed parameters
- Sessions belong to one user; ACLs must strictly prevent cross-user access
- Best practice: mask/remove PII before writing to storage, reducing blast radius and aiding GDPR/CCPA compliance
- Event append order must be deterministic (log disorder breaks reasoning)
- Sessions shouldn't be kept forever: use TTL or archival policies to control cost and compliance risk
- Generate summaries asynchronously in the background and persist them; avoid re-summarizing every turn
- Track which events are already covered by summaries to avoid re-injecting raw text
- Count-based: compress when turns/tokens exceed a threshold (most common, "good enough")
- Time-based: compress in the background after user inactivity
- Event-based: compress when a subtask/topic ends
- Personalization (preferences, facts, history)
- Replacing long histories (summaries/key facts reduce tokens)
- Insight at scale (aggregate trend mining, with privacy protection)
- Self-improvement (recording successful strategies/tool paths as playbooks)
- RAG: injects external, static, authoritative facts; generally shared, read-only
- Memory: accumulates user-specific, dynamic, isolated context; needs writing and evolution
- content: fact fragments (text or structured JSON/dict)
- metadata: id, owner, tags, source, etc.
- Declarative (knowing what): facts/preferences/events
- Procedural (knowing how): skills/processes/workflow "playbooks"
- Collections: piles of atomic memories, good for search and multi-topic
- Structured User Profile: stable fields (name, preferences), fast reads
- Rolling Summary: one continuously updated summary, often used to compress long sessions
- Vector databases: semantic similarity retrieval for unstructured memories
- Knowledge graphs: entity-relation reasoning for structured associations
- Hybrid: graph nodes with embeddings, supporting both semantic and relational retrieval
- Schema/template extraction (structured output)
- Natural language topic definitions
- Few-shot examples (especially effective for subtle domains)
- Duplicates (same fact, multiple expressions)
- Conflicts (changing user preferences)
- Evolution (facts becoming more specific)
- Decay and forgetting (old, low-confidence items get deleted/downweighted)
- Source type (preloaded system data / user input / tool output)
- Freshness (time)
- Confidence changes (rising with multi-source corroboration, decaying over time)
- Trust tiers during consolidation (trusted sources first, recent first, multi-source corroboration first)
- When deleting a data source, precisely regenerate via lineage instead of wiping everything
- Session end: cheap, but possibly low fidelity
- Every N turns: common middle ground
- Every turn, real-time: freshest, most expensive
- Explicit user instruction: controllable but incomplete coverage
- Memory-as-a-Tool: the Agent itself decides when to call
create_memory— smarter, but needs stricter tool definitions and policy controls to avoid "over-remembering" - Relevance (semantic)
- Recency (temporal freshness)
- Importance (annotated at generation time)
- Query rewriting (one extra LLM call)
- Reranking (take top 50, rerank with an LLM)
- Training a dedicated retriever (needs labeled data, costly) Pair with caching to avoid repeating expensive pipelines.
- Proactive prefetch: fetch at the start of every turn; simple but adds latency; cacheable
- Reactive (Memory-as-a-Tool): query only when needed; saves cost but may add an LLM call; the Agent must "know what memory types might exist," otherwise it won't know when to query
- Precision: how much recorded is correct and relevant (guards against "over-remembering" pollution)
- Recall: are key facts missed (guards against "critical fact loss")
- F1: combined metric
- Recall@K: does the needed memory appear in topK
- Latency: retrieval within a strict budget (e.g., <200ms) (Complex rewriting/reranking generally doesn't fit the real-time hot path unless cacheable and not easily stale.)
- Strict isolation (user/tenant scope): cross-user leakage is a fatal incident
- User control: opt-out and deletion supported
- PII masked before storage
- Defense against memory poisoning: validate and sanitize injected/forged information
- Cross-user shared procedural memory must be strongly anonymized, or "experience reuse" becomes "information leakage"
B. Evidence and factual data (the evidence chain)
C. Current interaction (the matter at hand)
> Tip: more context is not better. > Long context brings cost and latency, plus "context rot": the more information, the easier the model loses the key points and quality degrades. Context Engineering aims for "just enough, not too much."
Chapter 2: The Per-Turn "Context Pipeline" — Fetch → Prepare → Invoke → Upload
A production Agent typically loops through four steps each turn:
1. Fetch Context: retrieve needed context (memory, RAG, recent events) 2. Prepare Context: build the final prompt (this is the blocking hot path) 3. Invoke LLM and Tools: model and tool iterations append outputs to context 4. Upload Context: write new information to persistent storage (usually async in the background)
The most overlooked but critical point: Prepare is a blocking hot path; Upload should be backgrounded. Putting expensive distillation/consolidation work into the hot path makes the experience "lag on every message."
Chapter 3: Sessions — The Conversation Workbench (Event Log + Working State)
3.1 A Session consists of Events and State
A Session is a single-user, single-conversation container. Users can have multiple sessions, but they are disconnected "project workbenches" that don't naturally share memory (sharing requires Memory).
3.2 Session storage in production: runtimes are usually stateless
Most Agent runtimes are stateless: memory is cleared when a request ends. So session history must be written to persistent storage (database / managed session service) and pulled back at the start of each turn.Chapter 4: Framework Differences and Multi-Agent Session Patterns — Shared Big Ledger vs. Separate Small Ledgers
4.1 The essence of framework differences: unified internals, diverse external protocols
Frameworks are essentially "translators": developers use the framework's internal event structures, and the framework maps them to what each model API requires. This decouples model vendors but creates cross-framework "semantic isolation."4.2 Two patterns for organizing session history in multi-agent systems
The key is "how information is shared."#### A. Shared Unified History All agents write messages, tool calls, and observations into the same time-series log. Suitable for tightly coupled collaboration and pipeline tasks depending on a "single source of truth."
Pros: globally traceable, natural handoffs; Cons: history bloat, noise, harder permission isolation.
#### B. Separate Individual Histories Each agent has private history and only exposes final results (like a black-box tool), often via "Agent-as-a-Tool" or A2A messages exchanging results, not process.
Pros: clear boundaries, low leakage risk; Cons: thin shared context, collaboration needs extra design.
Chapter 5: The Hard Problem of Cross-Framework Interop — Sessions Aren't Portable, Memory Can Be the "Common Layer"
Different frameworks' session/event storage models are tightly bound to internal object structures, so a Session written by one framework is hard for another to read directly. A2A messages work, but "rich state" still needs a translation layer.
The more robust pattern: abstract shared knowledge into a framework-agnostic Memory layer. Memory stores not framework event objects but distilled facts/entities/summaries (strings or dicts), making it a shared cognitive resource across frameworks and agents.
Chapter 6: Production Considerations for Sessions — Isolation, Integrity, Performance, and Compression
6.1 Security and privacy: strict isolation + PII masking before storage
6.2 Data integrity: ordering consistency + lifecycle (TTL)
6.3 Performance: sessions sit on the hot path — must be "fast and small"
Each turn pulls session history to build the prompt; larger history means slower, costlier turns. Key optimization: filter/compress history before sending to the model (e.g., strip stale tool outputs).Chapter 7: Long Conversation Management — Packing History Like Luggage
Four hard constraints: 1) Context window limits 2) Token cost 3) Latency 4) Quality (noise and autoregressive error)
Compression strategies (simple to complex):
7.1 Sliding Window: keep the last N turns
Simple and effective, but may lose early key constraints.7.2 Token-Based Truncation: truncate by token budget
Better fits cost control; same risk of cutting key facts.7.3 Recursive Summarization: old content becomes summaries
Replace older segments with summaries coexisting with recent raw turns. Engineering points:7.4 Compression triggers: when?
Chapter 8: Memory — The Long-Term Filing Cabinet (Distilled, Reusable, "Less but Better")
Memory is not "storing the conversation" but extracting and solidifying valuable information that persists across sessions, used for personalization, context management, data insights, and self-improvement.
Four capabilities Memory provides
Chapter 9: Memory vs. RAG — One Knows the World, the Other Knows the User
One sentence: RAG makes the system a "facts expert"; Memory makes it a "user expert."
Chapter 10: Memory Structure and Types — Content + Metadata; "Knowing What" and "Knowing How"
10.1 Basic structure
10.2 Knowledge types: Declarative vs. Procedural
> Key difference: procedural memory is not "retrieving data" but "retrieving an executable plan" — more like a reasoning-augmentation layer. Compared with fine-tuning (offline, weight changes), procedural memory injects playbooks online, correcting behavior quickly via in-context learning.
Chapter 11: Organizing and Storing Memory — Collections / User Profiles / Rolling Summaries; Vector DBs / Knowledge Graphs / Hybrid
11.1 Organization patterns
11.2 Storage architectures
Chapter 12: Memory Generation — LLM-Driven ETL of Extraction + Consolidation
Memory generation isn't simple summarization but an LLM-driven ETL:
1) Ingestion: input source data (usually session history) 2) Extraction & Filtering: extract meaningful information per "topic definitions" (no match, no memory) 3) Consolidation: dedup, conflict resolution, update/create/delete, and forgetting (TTL/low confidence) 4) Storage: write to vector DBs/graphs and other persistence layers
12.1 Extraction: deciding "what's worth remembering"
"Meaningful" depends entirely on business goals. Common implementations:Many systems use "rolling summaries" as an auxiliary extraction input, improving efficiency and avoiding rescanning full history each turn.
12.2 Consolidation: without it, memory becomes a contradictory garbage heap
Consolidation must handle:Common operations: UPDATE / CREATE / DELETE (or INVALIDATE).
Chapter 13: Memory Provenance — Why Should You Trust It?
Long-term memory without source tracking drifts toward "confident nonsense." Provenance should at least record:
An important recommendation: generating long-term memory from tool outputs is generally not recommended — tool data suits short-term caching and is easily stale and fragile.
Engineering value of provenance:
Chapter 14: When to Generate Memory — Memory-as-a-Tool Is Smarter but Costs Must Be Controlled
Trigger strategy trade-off is "cost vs. fidelity":
Background vs. blocking: memory generation should almost always be async
Memory generation requires LLM calls and DB writes and must be stripped from the hot path: the user gets a reply first, extraction/consolidation happens in the background. Otherwise the experience becomes unusable.Chapter 15: Memory Retrieval — Relevance Isn't Enough; Recency and Importance Matter Too
The goal: find the most useful memories within a strict latency budget.
Common scoring dimensions:
Relying only on vector similarity is a common pitfall: it surfaces memories that are "similar but old/trivial." A multi-dimensional hybrid score is more robust.
Advanced enhancements (but slower):
Retrieval timing: proactive prefetch vs. reactive query
Chapter 16: Putting Memory into Context — Where You Put It Is Where the Weight Lands
Three main injection methods:
16.1 Into system instructions
Pros: high authority, clean conversation history; good for stable info (user profiles). Risks: over-influence — the model may force every topic toward the memory; system instructions are usually text-only (bad for multimodal memories); incompatible with "let the model decide whether to call a memory tool."16.2 Injected into conversation history
Pros: flexible; good for situational memories. Risks: noisy, high token cost, possible "conversation injection" illusion (the model treats memory as something the user just said). If injecting user-level memories with the user role, keep first-person perspective consistency.16.3 As tool output
With reactive querying, memories enter context as tool output, naturally landing in the conversation sequence. Pros: clear provenance; Cons: poor retrieval quality directly pollutes reasoning.Chapter 17: Evaluating Whether the Memory System Actually Works — Quality, Retrieval, End-to-End
Evaluation has three layers:
17.1 Generation quality: is it remembered correctly?
Compare against a human golden set:17.2 Retrieval performance: can it be found, and fast?
17.3 End-to-end task success: does memory improve task completion?
Use an LLM judge to compare final answers with golden answers, determining whether memory truly improves the target outcome.Chapter 18: Privacy and Security — Memory Is a "Corporate Archive," Not a "Casual Notepad"
The bottom line for long-term memory:
References (5 items)
1. Retrieval-Augmented Generation overview: https://cloud.google.com/use-cases/retrieval-augmented-generation?hl=en 2. Agent Engine Sessions overview: https://cloud.google.com/vertex-ai/generative-ai/docs/agent-engine/sessions/overview 3. Agent Engine Memory Bank – generate memories: https://cloud.google.com/vertex-ai/generative-ai/docs/agent-engine/memory-bank/generate-memories 4. Model Armor overview: https://cloud.google.com/security-command-center/docs/model-armor-overview 5. Gemini long context limitations: https://ai.google.dev/gemini-api/docs/long-context#long-context-limitations