Prompt Cache's technical principle is not complicated. During Transformer autoregressive generation, every generated token requires attention computation over all preceding tokens. Without caching, the K/V tensors for the entire history are recomputed at each step, giving O(n²) complexity. With KV Cache enabled, preceding tokens' K/V values are read directly from cache, and each new token only computes its own share, reducing complexity to O(n).
The technique has existed since the GPT-2 era, but between 2024 and 2026 it suddenly became an industry focus because it upgraded from an "inference optimization" into a commercial bottleneck.
The core insight of Prompt Cache: LLM input prompts contain large amounts of repeated text—system messages, prompt templates, context documents, few-shot examples. These segments are repeatedly fed into the model across many requests. Prompt Cache precomputes and stores the attention states (KV Cache) of these segments; when the same segment reappears, it is reused directly, and only the new portion is computed.
From KV Cache to Prompt Cache
Traditional KV Cache works only within a single request—K/V values of preceding tokens are cached for the next token. Prompt Cache extends this to the cross-request, cross-session level: when the same system prompt used by user A appears again for user B, its attention states are already computed and can be loaded directly.
Yale's paper "Prompt Cache: Modular Attention Reuse for Low-Latency Inference" provides a prototype implementation: latency reduced by 8x on CPU and 60x on GPU, with no modification to model parameters.
Key challenge: Transformers' positional encoding embeds each token's position into the attention states. If the same text appears at a different position, its attention states change and cannot be reused directly. The paper's solution is Prompt Markup Language (PML)—a schema that defines reusable "prompt modules," each assigned a unique position ID independent of global position. Experiments found that LLMs can handle attention states with non-contiguous position IDs: as long as the tokens' relative positions are preserved, output quality is unaffected.
Head Vendors' Prompt Cache Strategies in 2026
DeepSeek V4: Compression Is Justice
- 90% KV cache size compression via CSA (Cache-Strided Attention) / HCA (Hybrid Cache Attention) hybrid attention architecture
- 1M context standard across the lineup—1M is not top-tier, it's the baseline
- Cache hits priced at 10% of standard input token cost
- cache_control: explicit cache control at the API level; developers mark which parts are cacheable
- 4 cache breakpoints: prompts can be split into multiple cacheable segments, not just one prefix block
- Cross-session global sharing, 1-hour TTL: the same cache can be reused across different sessions
- Cache hit price = 1/10 of standard input
- Automatic caching: hidden behind the API; developers need no extra work
- Cache hit price = 50% of standard input
- No explicit cache control markers; hit rate is determined automatically by the system
- CachedContent API: an object-oriented abstraction for enterprise-scale long-material reuse
- 2M context window: currently the largest nominal context
- Explicit cache management: developers explicitly create, update, and delete cached content
DeepSeek's strategy is architecture-level compression + price leverage. Shrinking KV cache size to 1/10 turns long context from a "VRAM killer" into an "affordable option."
Anthropic Claude: Predictable Enterprise Contracts
Anthropic's strategy is predictability. Caching is not a black box but a promise written into the API contract. Boris Cherny, creator of Claude Code, acknowledged: "With the 1M context window, cache misses are very costly. If you leave your computer for over an hour and return to an old session, you usually won't hit the cache at all."
OpenAI GPT-5: Automatic Caching at 50% Off
OpenAI's strategy is ease of use first. Developers don't need to learn a new API—the system decides what can be cached. The cost is lower transparency and controllability.
Google Gemini: Object-Oriented Abstraction
Comparison Table
| Vendor | Cache Strategy | Cache Breakpoints | Hit Discount | TTL | Developer Control | |--------|----------------|-------------------|--------------|-----|-------------------| | DeepSeek | Architecture compression + automatic caching | Undisclosed | 10% | Undisclosed | Low | | Anthropic | Explicit cache_control | 4 | 10% | 1 hour | High | | OpenAI | Automatic caching | None | 50% | Undisclosed | Low | | Google | Object-oriented API | Undisclosed | Undisclosed | Developer-controlled | Very high |
---
> Key finding: Four vendors, four strategies, no unified standard. DeepSeek does architecture compression, Anthropic does contract predictability, OpenAI does ease of use, and Google does enterprise object management. This is itself an arms race of caching strategies.