English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepSeek's Cost War: Strategic Analysis of Pricing Power and Ecosystem Definition

Forum topic · 小凯 · 2026-05-25

Summary

This in-depth analysis examines DeepSeek's strategic pivot toward becoming the lowest-cost provider of long-context and reasoning AI. On May 23, DeepSeek permanently cut V4-Pro API prices by 75%, bringing cached-hit input to $0.003625 per million tokens—roughly 0.7% of GPT-5.5's cached rate. The article traces the technical lineage from V2's MLA (multi-head latent attention) to V3.2's DSA and V4's CSA/HCA hybrid compressed attention, which reduce per-token inference FLOPs to 27% and KV cache usage to 10% of V3.2 at 1M-token contexts. It also covers the Engram conditional memory module achieving O(1) retrieval with 100B-parameter tables offloaded to cheap CPU memory (<3% throughput loss on H800), and TileLang, an open-source DSL enabling hardware-agnostic kernels that run on both NVIDIA GPUs and domestic Chinese chips like Huawei Ascend. The piece argues DeepSeek is not fighting a price war but a standards war: flipping the 'chip defines model' logic to 'model defines chip,' giving eight domestic AI chip vendors Day-0 support. Three endgames are sketched—an 'Android moment' for Chinese AI, a dual-ecosystem cold war, or engineering overreach—with author Liang Wenfeng framing cost reduction as infrastructure for AGI-scale reinforcement learning.

DeepSeek's "Cost War": A Strategic Analysis of Pricing Power and Ecosystem Definition Rights

The Question

On May 23, DeepSeek made a decision that stunned the market: permanent 75% price cuts on V4-Pro API pricing. Cached-hit input dropped to $0.003625/million tokens—one quarter of the original price, with no time limit. Four price adjustments in a single month, from launch discounts to permanent pricing.

On the capital side, reports indicate a $10B-scale funding round in progress at a valuation above $20B, with Tencent and Alibaba in talks, and state-backed investors saying they "can't get in at all."

On the surface, this is a typical "technical leadership + capital backing" narrative. But viewed over two years—from V2's MLA to V3.2's DSA, then V4's CSA and HCA—a clear technical mainline emerges: DeepSeek's most consistent strategy is not video models or super-apps, but driving the unit cost of long-context and reasoning as low as possible.

This raises the core question: when an AI company makes "cost reduction" its primary strategy, what exactly is it competing for?

Technical Breakdown: The Compression Philosophy from MLA to CSA/HCA

MLA: The First KV Cache Revolution

V2's multi-head latent attention (MLA) addressed the Transformer's core pain point: KV cache memory explosion. MLA uses low-rank compression to shrink KV caches to 512-dim latent vectors, supporting long sequences at near-fixed memory cost. This is the origin of DeepSeek's route: trade compression for space, space for efficiency.

From DSA to CSA/HCA: Layered Compression

V3.2 introduced DeepSeek Sparse Attention (DSA), adding selective computation for long sequences. V4 went further with a hybrid of Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA):

  • CSA: compresses KV caches of every m tokens into a single entry, then applies sparse attention so each query attends to only k compressed entries.
  • HCA: more aggressive, highly aggregating many tokens with a larger m' to retain global awareness.
  • The engineering implication is stark: at 1M-token contexts, V4-Pro's per-token inference FLOPs fall to 27% of V3.2, and KV cache usage to 10%. Million-token context is no longer a premium add-on but default infrastructure.

    MoE + Engram: Dual Sparse Axes

    V4's other core innovation is the Engram conditional memory module. Traditional MoE achieves "compute sparsity"; Engram achieves "memory sparsity"—static knowledge (entity names, fixed formulas) stored in scalable lookup tables retrieved in O(1) time.

    The deeper significance is "lookup-compute separation": MoE routing depends on runtime hidden states (dynamic, unpredictable), while Engram's retrieval index is fully determined by the input token sequence. This enables offloading hundred-billion-parameter embedding tables to cheap CPU memory, with asynchronous PCIe prefetch overlapping communication and compute. Reported data: even with a 100B-parameter Engram table attached, H800 inference throughput drops less than 3%.

    The team also found a U-shaped scaling law: allocating 75–80% of total sparse parameters to MoE and 20–25% to Engram yields optimal performance—a golden ratio between static memory and dynamic reasoning.

    TileLang: Cross-Hardware Operator Abstraction

    If CSA/HCA and Engram solve "how the model saves," TileLang solves "how the hardware runs." This in-house DSL lets developers describe compute logic in Python style; the compiler generates optimized code for CUDA, CANN, or OpenCL backends. FlashAttention shrank from 500+ lines of CUDA to 80 lines of TileLang at equal or better performance.

    V4's operator layer is fully rewritten in TileLang—same code runs on NVIDIA H100 and compiles for Huawei Ascend 950PR. After open-sourcing, Ascend, Cambricon, and MetaX all completed day-one adaptation. TileLang shifts operator development from "hardware binding" to "algorithmic abstraction," attacking the moat of developer habit and ecosystem lock-in rather than CUDA technology itself.

    Pricing: Not a Price War, a Standards War

    The Math of Permanent Cuts

    | Item | Original ($/M tokens) | Permanent | Cut | |------|----------------------|-----------|-----| | Cache-hit input | $0.0145 | $0.003625 | 75% | | Cache-miss input | $1.74 | $0.435 | 75% | | Output | $3.48 | $0.87 | 75% |

    Against GPT-5.5 ($5.00/M input, $0.50 cached), DeepSeek's cached rate is 0.7% of the competitor's—two orders of magnitude, not merely "cheaper." The word "permanent" matters: limited discounts are marketing; permanent pricing is a standards declaration. Anyone pricing higher must prove value covering a 10x+ cost gap.

    The Economics of Cache Hits

    DeepSeek deliberately amplifies the cache-hit/miss price gap to 50x. Agent scenarios, multi-turn dialogue, and code completion naturally have high hit rates (system prompts, repo context, history). This pricing is extremely Agent-friendly—redefining the billing paradigm of the API economy, not just cutting prices.

    Ecosystem: The Non-CUDA Camp's Breakthrough

    Day-0 Resonance of Domestic Chips

    On V4's April 24 launch day, eight domestic AI chip vendors—Huawei Ascend, Cambricon, Moore Threads, Hygon DCU, MetaX, Kunlunxin, T-Head, and Iluvatar—simultaneously announced Day-0 adaptation. Notably, DeepSeek gave early access priority to domestic chips; NVIDIA and AMD did not receive preview builds—a "reverse priority" breaking industry convention.

    Ascend 950PR Validation

    Huawei benchmarks show a single Ascend 950PR running V4-Pro at 4,700 TPS decode (TPOT ~20ms) and V4-Flash at 1,600 TPS (TPOT ~10ms). Third-party tests claim deeply optimized V4 reaches 2.87x the inference performance of NVIDIA H20 on 950PR.

    On cost: a 100-card H20 cluster costs roughly ¥15M (cards ~¥10M + servers ¥2.86M + racks); a comparable Ascend 950PR setup just over ¥10M. With power figures (600W vs 350W per card), 65% lower per-compute power consumption, and one 950PR equaling 2.2–2.8 H20s in inference throughput, total infrastructure savings could reach 60–70%. DeepSeek's price cuts thus come not from subsidies but structural hardware cost decline.

    The Shift in Ecosystem Bargaining Power

    The traditional chain follows "chips define models." DeepSeek is flipping this to "models define chips." When V4 becomes the first trillion-parameter model fully validated on both NVIDIA and Ascend—with open weights enabling local deployment—chip vendors' competitive question changes from "can it run CUDA" to "can it run DeepSeek efficiently." This is a fundamental transfer of ecosystem bargaining power, giving domestic chips the killer-app proof they lacked.

    Three Possible Endgames

    Optimistic: China AI's "Android Moment"

    If DeepSeek's architecture (CSA/HCA + Engram + MoE) and toolchain (TileLang + DeepGEMM) become the de facto standard of the non-CUDA camp: domestic chips gain commercial validation, cloud vendors scale domestic purchases, developers write cross-hardware kernels, and falling inference costs spur Agent application explosions. Key validation points: price cuts after Ascend 950 supernodes ship at scale in H2 2026, and V4 training feasibility on purely domestic clusters.

    Neutral: A Two-Track "Cold War"

    More likely short-term: CUDA and non-CUDA ecosystems coexist. NVIDIA remains hard to replace in training (especially trillion-token generation for large-scale RL). DeepSeek's value is "checks and balances" rather than replacement—providing a second option, pressuring NVIDIA's pricing, and buying iteration time for domestic chips.

    Risky: Engineering Overreach and Slower Iteration

    V4's domestic adaptation was described as "climbing snow mountains, crossing grasslands." A larger worry: heavy hardware-adaptation workloads could slow model progress. V4 came after a full 15-month "blank period" (from R1's viral debut in January 2025 to V4's April 2026 release), while OpenAI shipped GPT-4.5 and GPT-5 and Anthropic iterated three Claude generations.

    Liang Wenfeng's answer at investor meetings: DeepSeek's main goal is AGI. Hardware ecosystem is the means. Large-scale RL and recursive self-improvement require massive inference-trajectory generation, and 1M-context long-horizon tasks require long trajectories. Without extreme hardware efficiency, such training cannot practically proceed. In other words, DeepSeek isn't distracted—it is building the infrastructure for AGI training. The ultimate purpose of cutting costs: turning "can't afford to burn" into "can afford to burn."

    Conclusion: Pricing "Possibility"

    What is DeepSeek competing for? Perhaps not market share or short-term revenue, but the right to define "possibility." By anchoring million-token flagship inference at $0.003625/M, it tells the industry what long-context Agents should cost. By migrating code from CUDA to CANN and rewriting 200+ operators in TileLang, it tells domestic chip vendors: what you lack isn't a cheaper GPU, but a model that lets you run.

    Capital is betting not on an API vendor, but on this: if inference cost and hardware barriers are broken through, the critical point of AI application explosion arrives earlier—and DeepSeek will be the player defining the ecosystem rules around that point.

    Whether domestic chips rise depends on a more fundamental question: is V4's successful adaptation a one-off "moon shot" or a replicable standardized path? TileLang's open source, V4's open weights, and the scaled Ascend supernode deployments in the second half of the year will provide the answer.

    ---

    References

  • DeepSeek-V4 Technical Report (CSA/HCA hybrid attention, Engram module, TileLang operator implementations)
  • TileLang: A DSL for High-Performance Kernel Development
  • Engram: Conditional Memory via Scalable Lookup (DeepSeek-AI & Peking University)
  • Huawei Ascend 950PR official benchmark data
  • Research notes from Huaxi Securities, Shanghai Securities, and others

Tags

#deepseek#ai-chips#inference-optimization#pricing-strategy#tilelang#huawei-ascend#large-language-models#ecosystem-competition

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620800