When you chat with Claude or GPT, every message causes the backend to re-encode your system prompt, tool definitions, conversation history, and new input from scratch. This prefill step is the main driver of latency and cost. After twenty turns, nineteen turns' worth of unchanged text gets recomputed every time. Prompt caching exists to end this waste.
How It Works: Prefix Matching
You mark a breakpoint in your request. The backend stores the encoding of everything up to that breakpoint. Next time, if the prefix is exactly identical, the cached encoding is reused and the duplicate computation is skipped.
The keyword is prefix matching—not "similar," not "close enough." Any change to a single character in the prefix invalidates all cache after that position.
The Numbers
- Cached tokens cost about 10% of normal input price.
- The first write costs 25% extra, but every subsequent hit saves 90%.
- Default TTL is 5 minutes, refreshed on each request.
- Minimum cacheable length: ~1024 tokens (4096 for newer models).
- Lazy loading: don't load all tools up front; group them by scenario and load on demand, keeping the prefix short and stable.
- Cache-safe forking: when compressing, summarizing, or branching a conversation, the forked request must share the same prefix—same model, system prompt, and tool definitions—or the cache resets to zero.
- Any change anywhere in the prefix invalidates everything after it.
- Use messages instead of instruction edits.
- Don't switch tools or models mid-conversation.
- Monitor cache hit rate like uptime.
- Forks must share the main conversation's prefix.
Example: a 100k-character conversation on Claude Sonnet costs ~$0.30 per turn without caching. With caching: $0.375 the first turn, then $0.03 per turn—roughly 90% savings on input cost over ten turns. More cache hits also mean lower TTFT (time to first token).
Anthropic monitors cache hit rate like server uptime: a drop triggers a paged incident with full incident-response procedures. Higher hit rates also let Anthropic offer paying users more generous usage limits. For Claude Code, caching isn't an optimization—it's the precondition for the product to work at all.
The Four-Layer Prompt Layout
Since caching is prefix-based, prompt ordering is critical. Arrange your prompt like a desk:
1. System instructions + tool definitions — fixed, shared across all sessions. 2. Project docs (CLAUDE.md) — shared within a project. 3. Current session context — valid for this conversation only. 4. Chat messages — grow turn by turn; only the last item is new.
Rule of thumb: the less something changes, the earlier it goes.
Three Common Pitfalls
1. Embedding the current time in the system prompt
Writing Today is {datetime.now()} in the system prompt changes the prefix on every request, killing the cache. Fix: keep the system prompt frozen and inject updated timestamps into the next message instead, e.g. wrapped in a <system-reminder> tag. System prompt = foundation (nailed down); messages = flowing water (change freely).
2. Unordered containers for tool definitions
Python dict or set can serialize tool definitions in different orders per request, breaking prefix matching. Fix: always use ordered lists.
3. Changing even one field of a tool
Adding a parameter or altering a field type in any tool invalidates the entire downstream cache, because tools sit in layer 1.
Counterintuitive Rule: Don't Switch Models
Switching from Opus to Haiku for easy questions seems sensible, but caches are model-bound. Switching models discards all accumulated cache, and rebuilding it often costs more than letting the big model answer the easy question directly. Claude Code therefore keeps one model for the main conversation throughout.
Need a smaller model? Delegate to sub-tasks. Sub-tasks get their own independent context and cache, so they don't pollute the main chain. The main model writes a task-handoff brief, the sub-task executes in isolation, and only the result returns. Claude Code's Explore agents work exactly this way.
Plan Mode: Change State via Tools, Not System Prompts
The intuitive approach to a planning mode would be removing execution tools when entering it—but that breaks the cache. Instead, Anthropic keeps the full tool set and adds two special tools: "enter planning" and "exit planning." The "no execution in planning" constraint is conveyed by inserting a system message into the conversation flow—not by modifying the system prompt. Bonus: the model can decide on its own when to plan.
Advanced Techniques
A Warning for API Resellers
Caches are isolated per account. Running a relay with an account pool yields hit rates too low to profit—and risks getting accounts banned. Similarly, avoid switching accounts mid-conversation.
The One Rule Behind All Seven
Every lesson here reduces to: caching is prefix matching.
---
*Based on the Prompt Cache interactive teaching module added in commit 515b759 of the easy-learn-ai project, which explains Prompt Cache design and best practices across 15 chapters in an animated-lecture format, inspired by Anthropic engineers' sharing on Claude Code prompt cache design.*