English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Prompt Cache Deep Dive: From Inference Optimization to Commercial Bottleneck

Forum topic · 小凯 · 2026-05-30

Summary

Prompt caching has evolved from a routine inference optimization into a commercial battleground for LLM providers. This post explains the technical foundation: autoregressive generation normally costs O(n²) without caching, while KV Cache reduces it to O(n). Prompt Cache extends this across requests and sessions by reusing attention states for repeated text segments such as system prompts, templates, and few-shot examples. Yale's paper "Prompt Cache: Modular Attention Reuse for Low-Latency Inference" achieved 8x latency reduction on CPU and 60x on GPU using Prompt Markup Language (PML) with modular position IDs. The post compares 2026 strategies: DeepSeek V4 compresses KV cache by 90% with strided/hybrid attention and charges 10% for cache hits with 1M context standard; Anthropic offers explicit cache_control with 4 breakpoints, 1-hour TTL, and 10% hit pricing; OpenAI GPT-5 uses automatic caching at 50% hit pricing; Google Gemini provides object-oriented CachedContent API with 2M context and developer-managed lifecycles. No unified standard exists across vendors.

Prompt Cache's technical principle is not complicated. During Transformer autoregressive generation, every generated token requires attention computation over all preceding tokens. Without caching, the K/V tensors for the entire history are recomputed at each step, giving O(n²) complexity. With KV Cache enabled, preceding tokens' K/V values are read directly from cache, and each new token only computes its own share, reducing complexity to O(n).

The technique has existed since the GPT-2 era, but between 2024 and 2026 it suddenly became an industry focus because it upgraded from an "inference optimization" into a commercial bottleneck.

The core insight of Prompt Cache: LLM input prompts contain large amounts of repeated text—system messages, prompt templates, context documents, few-shot examples. These segments are repeatedly fed into the model across many requests. Prompt Cache precomputes and stores the attention states (KV Cache) of these segments; when the same segment reappears, it is reused directly, and only the new portion is computed.

From KV Cache to Prompt Cache

Traditional KV Cache works only within a single request—K/V values of preceding tokens are cached for the next token. Prompt Cache extends this to the cross-request, cross-session level: when the same system prompt used by user A appears again for user B, its attention states are already computed and can be loaded directly.

Yale's paper "Prompt Cache: Modular Attention Reuse for Low-Latency Inference" provides a prototype implementation: latency reduced by 8x on CPU and 60x on GPU, with no modification to model parameters.

Key challenge: Transformers' positional encoding embeds each token's position into the attention states. If the same text appears at a different position, its attention states change and cannot be reused directly. The paper's solution is Prompt Markup Language (PML)—a schema that defines reusable "prompt modules," each assigned a unique position ID independent of global position. Experiments found that LLMs can handle attention states with non-contiguous position IDs: as long as the tokens' relative positions are preserved, output quality is unaffected.

Head Vendors' Prompt Cache Strategies in 2026

DeepSeek V4: Compression Is Justice

  • 90% KV cache size compression via CSA (Cache-Strided Attention) / HCA (Hybrid Cache Attention) hybrid attention architecture
  • 1M context standard across the lineup—1M is not top-tier, it's the baseline
  • Cache hits priced at 10% of standard input token cost
  • DeepSeek's strategy is architecture-level compression + price leverage. Shrinking KV cache size to 1/10 turns long context from a "VRAM killer" into an "affordable option."

    Anthropic Claude: Predictable Enterprise Contracts

  • cache_control: explicit cache control at the API level; developers mark which parts are cacheable
  • 4 cache breakpoints: prompts can be split into multiple cacheable segments, not just one prefix block
  • Cross-session global sharing, 1-hour TTL: the same cache can be reused across different sessions
  • Cache hit price = 1/10 of standard input
  • Anthropic's strategy is predictability. Caching is not a black box but a promise written into the API contract. Boris Cherny, creator of Claude Code, acknowledged: "With the 1M context window, cache misses are very costly. If you leave your computer for over an hour and return to an old session, you usually won't hit the cache at all."

    OpenAI GPT-5: Automatic Caching at 50% Off

  • Automatic caching: hidden behind the API; developers need no extra work
  • Cache hit price = 50% of standard input
  • No explicit cache control markers; hit rate is determined automatically by the system
  • OpenAI's strategy is ease of use first. Developers don't need to learn a new API—the system decides what can be cached. The cost is lower transparency and controllability.

    Google Gemini: Object-Oriented Abstraction

  • CachedContent API: an object-oriented abstraction for enterprise-scale long-material reuse
  • 2M context window: currently the largest nominal context
  • Explicit cache management: developers explicitly create, update, and delete cached content
Google's strategy is enterprise-grade management. CachedContent is a standalone API object with lifecycle management, suited for knowledge bases and document reuse in enterprise scenarios.

Comparison Table

| Vendor | Cache Strategy | Cache Breakpoints | Hit Discount | TTL | Developer Control | |--------|----------------|-------------------|--------------|-----|-------------------| | DeepSeek | Architecture compression + automatic caching | Undisclosed | 10% | Undisclosed | Low | | Anthropic | Explicit cache_control | 4 | 10% | 1 hour | High | | OpenAI | Automatic caching | None | 50% | Undisclosed | Low | | Google | Object-oriented API | Undisclosed | Undisclosed | Developer-controlled | Very high |

---

> Key finding: Four vendors, four strategies, no unified standard. DeepSeek does architecture compression, Anthropic does contract predictability, OpenAI does ease of use, and Google does enterprise object management. This is itself an arms race of caching strategies.

Tags

#prompt-cache#kv-cache#llm-inference#deepseek#anthropic-claude#openai-gpt-5#google-gemini#inference-optimization

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980604