China's LLM call volume tops US for 15 straight weeks

Share
China's LLM call volume tops US for 15 straight weeks

The week's usage data landed, and the curve keeps pointing one way: Chinese models now move more traffic than American ones — by a wide margin. The catch is in what the numbers actually measure.

Chinese AI models have now out-called their US rivals for 15 consecutive weeks, with the top four most-used models on the OpenRouter aggregation platform all coming from China. For the week of August 3–9, the platform recorded 69 trillion tokens of total usage, up 21.5% week over week. Chinese models accounted for 34.25 trillion of that — a 21.8% jump — against 9.17 trillion for US models, which grew 109% but from a smaller base. DeepSeek's freshly released V4-Flash led everything with 8.83 trillion tokens in a week (up 570%), followed by Tencent's Hy3, the DeepSeek preview build, and Xiaomi's MiMo-V2.5. OpenAI's GPT-5.6 Luna ranked fifth.

The numbers track a broader pricing collapse. Investment bank Jefferies, citing research firm Silicon Data, put average inference prices at $1.16–$1.18 per million tokens for August 6–8 — the year's low, down from $2.04 at the end of May. On OpenRouter, some Chinese models cost as little as 18 cents per million tokens against roughly $4 for US models, a more than 20x gap. Cheaper training, MoE architectures that activate a fraction of parameters per query, and low-cost power and data centers all feed the spread — and agent workloads, which burn tokens in bursts, reward it. US startups are noticing: agent platform Lindy says switching from Claude to DeepSeek-V4 cut its inference bill around 95%.

Read the caveats before calling a winner. OpenRouter only sees traffic routed through its gateway — third-party developers and downstream apps — not first-party apps, official websites, or private deployments, which is why the Financial Times and others caution against reading this as "China beats America." What it does show is a real shift in developer preference: the market is rewarding models that are good enough and dramatically cheaper, and Chinese labs are winning that trade on usage volume while the capability gap, by the UK AI Safety Institute's estimate, narrows to 4–7 months.


Insta360 is taking its thumb camera into AI hardware, pairing Alibaba's Qwen with Google's Gemini. The camera maker pushed an AI voice assistant called "Kira" to its GO Ultra pocket camera, with mainland China devices running Alibaba's Qwen models and Hong Kong, Macau, Taiwan, and overseas units using Gemini. Kira handles two-way live translation, photo-based questions, and voice queries, waking on a "Hey Kira" command or a long press of the shutter — on-device voiceprint and intent recognition, cloud-side answers. Insta360's founder says the goal is to make GO Ultra a leading portable personal-AI device, and the company is betting on a category IDC says is growing more than 350% year over year. Qwen, meanwhile, is stacking hardware wins: more than nine consumer brands — including vivo, Honor, Anker, and now Insta360 — have signed on to put Alibaba's models inside their devices.

What to watch: whether the next OpenRouter week keeps the streak alive — and whether US labs answer with price cuts of their own.

Does call volume tell us who's winning AI, or just who's cheaper? Tell us in the comments.

Sources: 21财经 · 观察者网 Guancha · 金融界 JRJ · 腾讯新闻·中国经营报 · IT之家