Zhipu's API revenue jumps 27x as China's AI labs prove the token economy
China's listed AI labs are finally producing numbers you can audit, and the morning's filings say the same thing: selling tokens is starting to work. Alongside that, two very different bets on where the next decade of AI data comes from — a wearable rig that records human experience, and a million-token model small enough to run on a phone.
Zhipu's first results since its Hong Kong listing show revenue up nearly 400%, and almost all of it now comes from selling tokens. The Beijing lab reported first-half revenue of 953.9 million yuan (about $142 million), up roughly 400% year on year — more than it booked in all of 2025. The striking number is underneath: its MaaS platform brought in 825 million yuan, about 86% of total revenue, a jump of more than 27x from 29.1 million yuan a year earlier. Token usage on the platform grew over 40x, paid daily active users rose 603%, and average API pricing climbed roughly 101%. The loss narrowed to 2 billion yuan from 2.4 billion, even as R&D spending rose 36.6% to 2.1 billion yuan.
The pricing figure is the one worth arguing about. Token volume growing 40x is a demand story, and a familiar one — but average price going up 101% while every major Chinese lab cuts prices is the opposite of the narrative. It suggests Zhipu isn't winning on being cheap; it's winning on coding and cybersecurity workloads where buyers pay for capability, and its GLM-5.3 has been pitched directly against Anthropic's Mythos 5 on code review and vulnerability discovery. That's a narrower, defensible business than a general-purpose price war. The caveat is scale: analysts expect 5 billion yuan for the full year, which is real money but still a rounding error next to the run rates U.S. labs are posting. We broke down the listing mechanics in Deep Dive — Zhipu and MiniMax just broke the "sell the unlock" rule in Hong Kong.
Ropedia's HOMIE Gen2 is a 380-gram headset built on a simple complaint: AI has read everything and experienced nothing. The Physical AI company, co-founded by CEO Chen Zhaoxi and CTO Hong Fangzhou with NTU associate professor Liu Ziwei as chief scientist, launched a second-generation multimodal capture rig that records what a person does rather than what the internet says. The device pairs four panoramic cameras with four spatial audio channels, funnels more than a dozen sensor streams onto one shared timeline and coordinate system with 50-microsecond sync, and needs no external setup — it runs for about 13 hours in homes, factories, or shops. Ropedia says deployment is roughly 10x faster than its previous approach and capture costs about one-twelfth as much.
The pitch is that video isn't experience, and neither is RGB, motion capture, or depth alone — only action, intent, environment, timing, and multiple senses aligned together qualify. Ropedia reports spatial tracking at roughly 0.2% relative error, 4.68 mm average error on 21-point hand tracking, under 2.5% depth error, and 96.0% action-labeling accuracy. Its Xperience-10M dataset already holds about 10 million real-world interaction clips and 10,000 hours of first-person video with audio, close to 1PB total. The company says it has served more than 20 robotics and foundation-model teams. Our read: the interesting claim here is the industrialization thesis — Liu argues that if experience is the scarce resource, someone has to manufacture it, and information lost at capture can't be recovered by any algorithm downstream. That's a real argument, and it's also the argument of a company selling the shovels.
iFlytek open-sourced two small Spark X2.5 models with a 1-million-token context window — the first edge models to carry one. The 4B and 1.7B versions landed September 1 with gains in agent interaction, math, and general reasoning, aimed at in-car systems, smart hardware, and IoT. A 293B Spark X2.5 base model follows on September 7, with upgrades to code generation and agent collaboration.
A million-token context used to be a datacenter flex. Putting it in a 1.7B model reframes what on-device AI can hold: not a chat history, but an entire working context — a vehicle's service record, a sensor's full day, a codebase — resident locally, with no round trip. The practical wins are latency, privacy, and offline operation; the open question is whether a model that small actually uses a window that large, or whether the number mostly demos well. Worth watching what builders ship with it.
What to watch: Zhipu's September 7 base-model release from iFlytek, and whether other Chinese labs start reporting API-line revenue this cleanly — that's the number that decides if this is a durable business or a good half-year.
If AI's next bottleneck is experience rather than text, does it matter that the training data can only be gathered by putting cameras on people? Tell us in the comments.
Sources: Reuters · 雷峰网 Leiphone · Business Wire · AI Base · Gate News