China's labs raise API prices while third-party hosts slash them
Two price curves for the same open weights are now moving in opposite directions — and that split says more about where open-model economics are heading than any benchmark released this week.
DeepSeek and Zhipu have hiked their official API prices, while third-party hosts serving the exact same open weights have cut theirs — sometimes by half. A 21st Century Business Herald analysis lays out the contradiction in detail: DeepSeek's August repricing raised peak output to 27 yuan per million tokens, cache-miss input to 9 yuan, and cache-hit input to 0.30 yuan — roughly 4.5x, 3x and 12x their previous levels, with off-peak set at half of peak. Zhipu's GLM Coding Plan tiers jumped even harder, with Lite, Pro and Max monthly subscriptions rising about 141%, 261% and 130% while the billing model shifted from request caps to token-based credits. Yet on OpenRouter, dozens of providers are serving the same DeepSeek V4-Flash-0731 for as little as $0.05 per million input tokens — roughly a quarter of DeepSeek's own off-peak rate. We covered the August hike when it landed — DeepSeek hikes API prices up to 4.7x with new peak-hour billing — and framed it as the price war ending. That reading now looks incomplete: the war relocated to someone else's infrastructure.
The reason the two curves diverge is that a lab and a host are not selling the same thing. DeepSeek carries training, post-training, safety evaluation, version iteration and the cost of reserving inference capacity for peak concurrency — a coding agent that resends its system prompt, tool definitions and repo history on every turn leans hard on context caching, and cached KV state keeps occupying GPU memory even when it saves recompute. That is why the steepest increase landed on cache-hit input, the line item agent traffic depends on most. A host, by contrast, only has to run inference cheaply: cheaper or domestic accelerators, lower-precision quantization, smarter batching and scheduling, or simply promotional pricing to buy utilization. As Song Fei, founder of Singapore-based Infinite Alignment, puts it in the piece, the labs are asserting pricing power over their own endpoint while the hosts are competing on the marginal cost of running weights they got for free.
The interesting counter-move is Moonshot's, and it isn't a price at all. Kimi K3 shipped under a custom license rather than MIT, requiring any business earning more than $20 million across its affiliates selling inference or fine-tuning as a service to negotiate separately — plus a branding clause that forces products above 100 million monthly users or $20 million monthly revenue to display "Kimi K3" in the interface. The weights are still downloadable, modifiable and commercially usable, so nothing changes for most developers. But K3 released at low training precision, which closes off most of the quantization headroom a host would normally exploit, and Moonshot keeps the serving-side knowledge — scheduling, cache policy, hot prefixes — that only comes from running your own traffic. We covered the revenue-share play when it surfaced — Moonshot wants a cut when US clouds sell its Kimi K3 model — and this is the other half of the same strategy.
The distribution angle is the one to watch. Tmall launched an "AI Space Station" token top-up center on September 3, putting subscription plans from Alibaba Cloud, Zhipu, Kimi and MiniMax on a standard e-commerce shelf with card-key or direct top-up delivery, after Zhipu opened its own flagship store a day earlier and reportedly saw search volume rise 40-fold on day one. Alibaba's listings there undercut its own site by one cent. When model subscriptions sit next to phone plans and get compared on price, the vendor's own storefront stops being the default channel — and the labs that can't win on token price have to win on workflow lock-in instead.
What to watch: whether Moonshot's license approach spreads. If more labs start shipping weights in forms that resist cheap rehosting, "open weights" quietly becomes a distribution channel rather than a pricing floor.
If your model bill is up but your per-token rate is down, which provider are you actually routing to — and do you know? Tell us in the comments.
Sources: 21st Century Business Herald / Southern Finance · Pandaily · Implicator.ai · OpenRouter — DeepSeek V4 Flash 0731 · Jiemian via Tencent News · Caiwen via Tencent News