English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepSeek V4 Peak/Off-Peak API Pricing Takes Effect: 1100% Cache Hit Hike Forces Budget Reworks

Forum topic · 小凯 · 2026-08-17

Summary

Starting August 17 at midnight, DeepSeek V4 series APIs adopted time-of-use pricing modeled on electricity peak-valley schemes. Peak hours (Beijing time 9:00-12:00, 14:00-18:00) cost twice the off-peak rate. The steepest increase hits cache-hit input: V4 Pro peak price is 0.3 yuan per million tokens, a 1100% jump from 0.025 yuan, while cache-miss input rose 200% (9 yuan) and output 350% (27 yuan). V4 Flash saw smaller hikes (400% cache-hit). The move follows a V4-Flash single-day record of ~8 trillion tokens on August 1 and an API outage on August 4. DeepSeek is the first Chinese LLM vendor to bring peak-valley scheduling to token-level API pricing, unlike Anthropic, OpenAI, and Google, which use direct hikes or subscription tiers. Developers running cache-heavy workloads—agents, multi-turn chat, document review, code review—face disproportionate cost growth and are advised to shift batch jobs to off-peak hours (half price) or fall back to alternatives like Qwen3.8-27B or GLM-5.3 during peaks.

DeepSeek has applied the "electricity peak-valley pricing" model to its LLM API.

Starting August 17 at 00:00, the DeepSeek V4 series APIs officially adopted an industry-rare time-of-use (peak/off-peak) pricing scheme. Peak hours (Beijing time 9:00-12:00 and 14:00-18:00) are priced at 2x the idle-hour rate. This is DeepSeek's fourth pricing adjustment announcement since the start of 2026.

The Numbers

The most eye-catching hike is on cache-hit input:

  • V4 Pro, peak hours: cache-hit input 0.3 yuan / million tokens, up 1100% from the original 0.025 yuan; cache-miss input 9 yuan (up 200%); output 27 yuan (up 350%). Off-peak: 4.5 / 13.5 / 0.15 yuan respectively.
  • V4 Flash, peak hours: cache-hit input 0.1 yuan (up 400%); cache-miss input 3 yuan (up 200%); output 9 yuan (up 350%). Off-peak: 0.05 / 1.5 / 4.5 yuan respectively.
  • Why DeepSeek Can Do This

    The company's confidence comes from load pressure. Two dates are frequently cited: on August 1, V4-Flash processed roughly 8 trillion tokens in a single day; on August 4, the API briefly went unavailable due to traffic. This gives "off-peak at half price" an economic rationale—developers who move workloads to late night cut their bills in half—and "peak surcharge" a compute rationale: scarce peak-hour capacity goes to latency-sensitive, premium-paying workloads.

    The Real Impact: Cache-Heavy Workloads

    The real damage lands on the cache-hit line, and the 1100% figure understates the effect. Businesses with high cache-hit ratios see cost growth far exceeding the headline percentage. Anything relying on large repeated prefixes—agents, multi-turn conversations, document review, code review—will see bills reworked from scratch. DeepSeek is effectively using price signals to redraw the boundary between "cache-heavy" and "generation-heavy" businesses: the former are pushed toward idle hours, the latter toward paying peak premiums.

    Why It Matters

    This is the first time a Chinese LLM vendor has brought grid-dispatch thinking to LLM API monetization. The previous domestic playbook was: low prices to acquire users → scale → raise prices. DeepSeek teased higher pricing in its API docs on August 6, quietly raised V4-Pro-0813's cache-hit price from 0.025 to 0.1 yuan per million tokens on August 12, formally announced the time-of-use scheme on August 13, and executed on August 17—a four-step cadence more precise than any OpenAI price change.

    For comparison, Anthropic, OpenAI, and Google Gemini do not offer time-of-use pricing. They either raise prices directly (GPT-5.6 cache-hit input at $0.05/million tokens, about 1.25x DeepSeek's post-hike $0.04) or differentiate via subscription tiers (e.g., Claude Max at $200/$400). DeepSeek's approach resembles the Spot Instance path used by Alibaba Cloud and Tencent Cloud for inference compute—but this is the first time it has been applied at the token level.

    What Developers Should Do

  • Cache-heavy workloads: re-measure your bills now; move offline batch processing out of real-time paths
  • Peak-shifting: push night jobs, batch tasks, and deferrable tasks into idle hours for an immediate 50% discount
  • Multi-model strategy: consider falling back to domestic alternatives like Qwen3.8-27B or Zhipu GLM-5.3 during peak hours
  • What It Means for the Industry

  • The "subsidize-for-scale" phase is over; DeepSeek V4 enters a "scale-for-profit" phase, with steep revenue elasticity as daily volume climbs beyond 8 trillion tokens
  • Chinese LLMs are shifting from "competing on price" to "competing on pricing mechanisms": time-of-use, tiered, subscription, token packs, and enterprise contracts will run in parallel
Three things to watch: DeepSeek V4's actual peak-hour margins over the next 12 months; whether developer goodwill holds in idle hours; and whether other domestic vendors (Qwen / Zhipu / Kimi / ByteDance Doubao) follow suit. If more than two vendors adopt time-of-use pricing within three months, the Chinese LLM API market will enter its own "electricity peak-valley" era.

Tags

#deepseek#api-pricing#llm#time-of-use-pricing#cache-hit#cost-optimization#china-ai#v4

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633580