DeepSeek has applied the "electricity peak-valley pricing" model to its LLM API.
Starting August 17 at 00:00, the DeepSeek V4 series APIs officially adopted an industry-rare time-of-use (peak/off-peak) pricing scheme. Peak hours (Beijing time 9:00-12:00 and 14:00-18:00) are priced at 2x the idle-hour rate. This is DeepSeek's fourth pricing adjustment announcement since the start of 2026.
The Numbers
The most eye-catching hike is on cache-hit input:
- V4 Pro, peak hours: cache-hit input 0.3 yuan / million tokens, up 1100% from the original 0.025 yuan; cache-miss input 9 yuan (up 200%); output 27 yuan (up 350%). Off-peak: 4.5 / 13.5 / 0.15 yuan respectively.
- V4 Flash, peak hours: cache-hit input 0.1 yuan (up 400%); cache-miss input 3 yuan (up 200%); output 9 yuan (up 350%). Off-peak: 0.05 / 1.5 / 4.5 yuan respectively.
- Cache-heavy workloads: re-measure your bills now; move offline batch processing out of real-time paths
- Peak-shifting: push night jobs, batch tasks, and deferrable tasks into idle hours for an immediate 50% discount
- Multi-model strategy: consider falling back to domestic alternatives like Qwen3.8-27B or Zhipu GLM-5.3 during peak hours
- The "subsidize-for-scale" phase is over; DeepSeek V4 enters a "scale-for-profit" phase, with steep revenue elasticity as daily volume climbs beyond 8 trillion tokens
- Chinese LLMs are shifting from "competing on price" to "competing on pricing mechanisms": time-of-use, tiered, subscription, token packs, and enterprise contracts will run in parallel
Why DeepSeek Can Do This
The company's confidence comes from load pressure. Two dates are frequently cited: on August 1, V4-Flash processed roughly 8 trillion tokens in a single day; on August 4, the API briefly went unavailable due to traffic. This gives "off-peak at half price" an economic rationale—developers who move workloads to late night cut their bills in half—and "peak surcharge" a compute rationale: scarce peak-hour capacity goes to latency-sensitive, premium-paying workloads.
The Real Impact: Cache-Heavy Workloads
The real damage lands on the cache-hit line, and the 1100% figure understates the effect. Businesses with high cache-hit ratios see cost growth far exceeding the headline percentage. Anything relying on large repeated prefixes—agents, multi-turn conversations, document review, code review—will see bills reworked from scratch. DeepSeek is effectively using price signals to redraw the boundary between "cache-heavy" and "generation-heavy" businesses: the former are pushed toward idle hours, the latter toward paying peak premiums.
Why It Matters
This is the first time a Chinese LLM vendor has brought grid-dispatch thinking to LLM API monetization. The previous domestic playbook was: low prices to acquire users → scale → raise prices. DeepSeek teased higher pricing in its API docs on August 6, quietly raised V4-Pro-0813's cache-hit price from 0.025 to 0.1 yuan per million tokens on August 12, formally announced the time-of-use scheme on August 13, and executed on August 17—a four-step cadence more precise than any OpenAI price change.
For comparison, Anthropic, OpenAI, and Google Gemini do not offer time-of-use pricing. They either raise prices directly (GPT-5.6 cache-hit input at $0.05/million tokens, about 1.25x DeepSeek's post-hike $0.04) or differentiate via subscription tiers (e.g., Claude Max at $200/$400). DeepSeek's approach resembles the Spot Instance path used by Alibaba Cloud and Tencent Cloud for inference compute—but this is the first time it has been applied at the token level.