Zhipu's GLM-5.3-FlashX hits 200 tokens/s — at 2.5x the price
Zhipu shipped the fast tier of its open model today and, unusually for a Chinese lab in 2026, raised the price to do it.
GLM-5.3-FlashX is live on the API at up to 200 tokens per second — roughly five times the throughput of GLM-5.3-Flash — and priced at about 2.5x what the base model costs. Zhipu says its 100,000-chip domestic inference cluster was saturated at launch under the weight of demand that never let up, and that the new tier follows fresh infrastructure investment and inference-side optimization rather than any change to the underlying model. It runs under its own model key, so developers pay the speed premium per call instead of inheriting it.
The base is the model we know. GLM-5.3-Flash shipped and open-sourced in August with 320 billion total parameters and 18 billion active in a mixture-of-experts layout — the first natively multimodal entry in the GLM-5 line, aimed at coding, agent work, visual programming and document handling. It also ran as the anonymous "Ox Alpha" on OpenRouter and OpenCode, where it became the highest-volume model on both before anyone knew whose it was. We covered the benchmark side of its launch in August — Z.ai's GLM-5.3-Flash tops benchmarks at one-tenth the price.
The interesting part is the pricing direction. For two years the Chinese open-model story was a race to zero, and a "Flash" badge meant cheap. FlashX keeps the same intelligence and the same architecture — hybrid linear-plus-sparse attention, fewer active parameters, native vision for checking its own front-end and Office output — and charges more for latency instead. That is a compute-scarcity admission: the constraint on serving agent traffic in China is no longer model quality, it is tokens per second on domestic silicon, and the lab is now selling the number customers can't get elsewhere. It also carves the market in two — price-sensitive batch work stays on Flash, latency-sensitive coding agents get billed a premium for waiting less.
Huawei used its Shanghai conference today to put a name on the same problem and a date on its export answer. At Huawei Connect 2026, Huawei Cloud CEO Peter Zhou said the company's enterprise agent platform AgentArts — already running inside more than 100 organizations, among them the Shenzhen Longgang district government, China Southern Power Grid and Kingsoft Office — will be commercially available outside China on December 30. Its open-source edition, openJiuwen, now carries more than 5,000 general and 1,000 industry-specific MCP assets, 50,000 GitHub stars and 3.29 million downloads, and Huawei is packaging partner commercial editions with Chinasoft International, iSoftStone and Beiming Software.
Huawei also published a seven-step "AI landing" method and the DIMAK engineering system spanning data, infrastructure, models, agents and knowledge, pitched at the failure mode of scattered agents with no process owner. Rotating chairman Eric Xu separately introduced the Peerium computing architecture, which links processors, NPUs, memory, storage and switches over an open bus protocol — the Atlas 950 supernode is its first commercial generation. The through-line is that Chinese infrastructure vendors are now selling deployment methodology, not just silicon.
What to watch: whether the speed premium holds once rivals push similar tiers, and whether AgentArts clears Western procurement review when it lands outside China.
Would you pay 2.5x for a faster model, or route the slow one and wait? Tell us in the comments.
Sources: AIHub · OSCHINA · Sina Tech via Sohu · AIBase · Huawei · ChinaTechNews · 智东西 Zhidx