English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

How AI Takes Shortcuts: The Counterintuitive Survival Rules Behind Prompt Caching

Forum topic · 小凯 · 2026-05-23

Summary

Prompt caching eliminates redundant computation in LLM inference by reusing the encoded prefix of a request when it matches exactly across calls. This article explains why caching matters: without it, every conversation turn re-encodes the full system prompt, tool definitions, and history, inflating both cost and time-to-first-token. Cached tokens cost roughly 90% less after a 25% one-time write premium, and Anthropic treats cache hit rate as infrastructure-grade telemetry. The piece distills best practices: order prompts in fixed layers (system instructions and tools first, messages last), keep tool definitions in stable ordered containers, inject dynamic content like timestamps into messages rather than system prompts, never switch models mid-conversation (caches are model-bound—delegate to sub-agents instead), and implement Plan Mode via special tools plus in-flow system messages rather than tool-set changes. It also covers lazy tool loading and cache-safe forking, and warns API resellers that caches are account-isolated.

When you chat with Claude or GPT, every message causes the backend to re-encode your system prompt, tool definitions, conversation history, and new input from scratch. This prefill step is the main driver of latency and cost. After twenty turns, nineteen turns' worth of unchanged text gets recomputed every time. Prompt caching exists to end this waste.

How It Works: Prefix Matching

You mark a breakpoint in your request. The backend stores the encoding of everything up to that breakpoint. Next time, if the prefix is exactly identical, the cached encoding is reused and the duplicate computation is skipped.

The keyword is prefix matching—not "similar," not "close enough." Any change to a single character in the prefix invalidates all cache after that position.

The Numbers

  • Cached tokens cost about 10% of normal input price.
  • The first write costs 25% extra, but every subsequent hit saves 90%.
  • Default TTL is 5 minutes, refreshed on each request.
  • Minimum cacheable length: ~1024 tokens (4096 for newer models).
  • Example: a 100k-character conversation on Claude Sonnet costs ~$0.30 per turn without caching. With caching: $0.375 the first turn, then $0.03 per turn—roughly 90% savings on input cost over ten turns. More cache hits also mean lower TTFT (time to first token).

    Anthropic monitors cache hit rate like server uptime: a drop triggers a paged incident with full incident-response procedures. Higher hit rates also let Anthropic offer paying users more generous usage limits. For Claude Code, caching isn't an optimization—it's the precondition for the product to work at all.

    The Four-Layer Prompt Layout

    Since caching is prefix-based, prompt ordering is critical. Arrange your prompt like a desk:

    1. System instructions + tool definitions — fixed, shared across all sessions. 2. Project docs (CLAUDE.md) — shared within a project. 3. Current session context — valid for this conversation only. 4. Chat messages — grow turn by turn; only the last item is new.

    Rule of thumb: the less something changes, the earlier it goes.

    Three Common Pitfalls

    1. Embedding the current time in the system prompt

    Writing Today is {datetime.now()} in the system prompt changes the prefix on every request, killing the cache. Fix: keep the system prompt frozen and inject updated timestamps into the next message instead, e.g. wrapped in a <system-reminder> tag. System prompt = foundation (nailed down); messages = flowing water (change freely).

    2. Unordered containers for tool definitions

    Python dict or set can serialize tool definitions in different orders per request, breaking prefix matching. Fix: always use ordered lists.

    3. Changing even one field of a tool

    Adding a parameter or altering a field type in any tool invalidates the entire downstream cache, because tools sit in layer 1.

    Counterintuitive Rule: Don't Switch Models

    Switching from Opus to Haiku for easy questions seems sensible, but caches are model-bound. Switching models discards all accumulated cache, and rebuilding it often costs more than letting the big model answer the easy question directly. Claude Code therefore keeps one model for the main conversation throughout.

    Need a smaller model? Delegate to sub-tasks. Sub-tasks get their own independent context and cache, so they don't pollute the main chain. The main model writes a task-handoff brief, the sub-task executes in isolation, and only the result returns. Claude Code's Explore agents work exactly this way.

    Plan Mode: Change State via Tools, Not System Prompts

    The intuitive approach to a planning mode would be removing execution tools when entering it—but that breaks the cache. Instead, Anthropic keeps the full tool set and adds two special tools: "enter planning" and "exit planning." The "no execution in planning" constraint is conveyed by inserting a system message into the conversation flow—not by modifying the system prompt. Bonus: the model can decide on its own when to plan.

    Advanced Techniques

  • Lazy loading: don't load all tools up front; group them by scenario and load on demand, keeping the prefix short and stable.
  • Cache-safe forking: when compressing, summarizing, or branching a conversation, the forked request must share the same prefix—same model, system prompt, and tool definitions—or the cache resets to zero.
  • A Warning for API Resellers

    Caches are isolated per account. Running a relay with an account pool yields hit rates too low to profit—and risks getting accounts banned. Similarly, avoid switching accounts mid-conversation.

    The One Rule Behind All Seven

    Every lesson here reduces to: caching is prefix matching.

  • Any change anywhere in the prefix invalidates everything after it.
  • Use messages instead of instruction edits.
  • Don't switch tools or models mid-conversation.
  • Monitor cache hit rate like uptime.
  • Forks must share the main conversation's prefix.
It looks like cache optimization, but it's also a design philosophy: accept the unavoidable constraint first, then build the entire system around it.

---

*Based on the Prompt Cache interactive teaching module added in commit 515b759 of the easy-learn-ai project, which explains Prompt Cache design and best practices across 15 chapters in an animated-lecture format, inspired by Anthropic engineers' sharing on Claude Code prompt cache design.*

Tags

#prompt-caching#anthropic#claude-code#llm#inference-cost#ai-agents#prompt-engineering#caching

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620686