Z.ai's GLM-5.3-Flash tops benchmarks at one-tenth the price

Share
Z.ai's GLM-5.3-Flash tops benchmarks at one-tenth the price

The mystery model that ruled the leaderboards all week finally showed its face today — and the reveal came with a price tag that's reshaping the frontier.

Z.ai has confirmed that GLM-5.3-Flash — the model it released today — is the same engine that topped benchmarks anonymously as "Ox Alpha" this week. The first natively multimodal model in the GLM-5 series packs 320 billion total parameters with just 18 billion active, a sparse MoE design the company says lets it beat its predecessor GLM-5.2 across the board at one-tenth of the price. On the Artificial Analysis Intelligence Index, the model scores 57 at roughly $0.045 a task — a level of intelligence that previously cost about ten times more.

On coding and agentic work it gets close to Anthropic's Claude Opus 4.8, according to Z.ai's own evaluation: 63.4 versus 46.2 on DeepSWE v1.1, 48.8 versus 26.2 on AutomationBench, and near-parity with Opus 4.8 on the lab's Z.ai Code Bench at max effort (29.0 to 29.5). The efficiency is architectural rather than cosmetic. Z.ai says a hybrid of linear and sparse attention, combined with a 30-trillion-token multimodal pre-training corpus, lets the model produce more intelligence with less compute — cutting attention cost and KV-cache footprint by roughly three- and four-fold versus GLM-5.3.

That same engine had been serving invisibly as "ox-alpha" on OpenCode and OpenRouter, where it became the most popular model of the week — settling the attribution question we tracked from its launch, in Mystery model Ox Alpha tops GPT-5.6 — sleuths say it's Zhipu's GLM. The reveal also answers where it ran: Z.ai says all of that hidden traffic was served on a large cluster of Chinese AI chips, and that after tuning an inference stack around them, per-token cost on that domestically built hardware is now roughly comparable to mainstream Nvidia GPUs. Weights are public under an MIT license, and the model already ships to GLM Coding Plan and ZCode users.

The deeper signal is the price collapse working its way up the frontier. A model approaching Claude Opus 4.8's coding ability at a tenth of the cost, trained and served on chips built under export controls, is exactly the trade-down pressure that's been squeezing OpenAI and Anthropic all year. What to watch: whether GLM-5.3-Flash's sub-$0.05 headline price holds under real traffic, and how quickly flagship pricing responds.

If a frontier-class coder now runs comfortably within a $0.05 budget, how long before coding-model pricing follows it down?

Sources: Z.ai announcement · TechCrunch · Techmeme · Reddit r/LocalLLaMA discussion