DeepSeek hikes API prices up to 4.7x with new peak-hour billing
The price war's last holdout just posted its new rate card. DeepSeek has turned its promised "significant" API increase into hard numbers — and, for the first time, into a peak/off-peak billing model that charges more than double for busy-hour inference.
DeepSeek has finalized its first major API price increase, with output rates rising roughly 2.4x to 4.7x and a new peak/off-peak billing structure taking effect at 16:00 UTC on August 16. Under the new schedule, a million output tokens from V4-Flash cost $0.66 off-peak and $1.32 at peak — up from $0.28 today — while flagship V4-Pro output goes from $0.87 to $1.98 off-peak and $3.96 at peak. Peak hours run 01:00–04:00 and 06:00–10:00 UTC (09:00–12:00 and 14:00–18:00 in Beijing), with off-peak rates set at exactly half of peak. Input tokens rise too: cache-miss input roughly 1.5x–3x, and cache-hit input as much as 5x on V4-Flash and 12x on V4-Pro — though the latter still lands at just 4.4 cents per million tokens.
The move matters because DeepSeek has been the price floor of the entire model market: V4-Flash topped OpenRouter's weekly volume chart at 8.83 trillion tokens, and its rates were the benchmark every competitor undercut. That floor just moved up, and the direction is unmistakable — the two-year price war is ending. The increase caps a week in which 01.AI began winding down its open model platform, Alibaba weighed revenue-sharing terms for Qwen's open weights, and the Qwen app launched paid memberships, a shift we flagged as it happened. DeepSeek held out longest, and the notice it posted on August 6 — which we noted alongside DeepSeek ships V4 Pro, ending its flagship's four-month preview yesterday — now has real numbers behind it.
The peak/off-peak mechanism is the quietly interesting part: DeepSeek is pricing its API the way utilities price electricity, using rates to push cost-sensitive workloads into off-peak windows and ration capacity at the exact hours its infrastructure is most loaded. For developers, the calculus changes from "pick the cheapest frontier model" to "pick the cheapest model at the hour you actually run it" — and for the rest of the industry, DeepSeek just handed every rival permission to raise prices, too.
What to watch: whether Qwen, Zhipu or Kimi move their own rates before August 16, and whether the increase finally dents the DeepSeek volumes that have led OpenRouter for fifteen straight weeks.
Did DeepSeek's price war really end — and what will a peak-hour surcharge do to your API bill? Tell us in the comments.
Sources: DeepSeek API pricing docs · 比特网 (Bitnet) · TMTPost