Zhipu's GLM-5.3 ships with 'emergent' cyber capabilities
China's open-weights race just produced its most awkward win: Zhipu's GLM-5.3 landed Thursday with frontier-level coding and a cyber capability its own makers admit outgrew their expectations — and Moonshot AI is racing a closing window toward a Hong Kong listing.
Zhipu released GLM-5.3 on Thursday, an open-weights flagship whose every gain comes from post-training — including an "emergent" cyber capability the company says developed faster than it expected. The model shares GLM-5.2's base; all the improvement comes from scaling reinforcement learning on long-horizon task environments. Zhipu reports a 50 percent jump over GLM-5.2 on its in-house Z.ai Code Bench, open-weights state of the art on Terminal Bench 3.0 (28.3 versus 4.6) and Agents' Last Exam, and — at High effort — better agentic coding results than Claude Opus 4.8 while using under half the output tokens. We flagged this launch back in August — Zhipu confirms GLM-5.3 launch as Goldman raises outlook 35% — and the release beats the preview.
The cyber numbers are the part that reads like a policy memo, not a launch note. Zhipu says GLM-5.3 tops CyberGym for vulnerability discovery (84.5 versus GLM-5.2's 77.2) and more than doubles its predecessor on exploitation benchmarks — 54.4 on ExploitBench versus 24.4, and 105 exploitation tasks completed within two hours versus 29. In real-world testing alongside Chinese security teams, the model found 2,436 vulnerabilities across 269 open-source projects, 1,097 of medium-to-high severity; the oldest flaw dates to 1981 and averaged 26.6 years undiscovered. Zhipu is tracking disclosures publicly via a security ledger, with 53 findings public and 2,383 under embargo, and will release weights in about two weeks, after safety evaluation and hardening. The honest framing in Zhipu's own post: capability is growing fastest exactly where open models trail furthest — the closed frontier (Anthropic's Mythos 5, OpenAI's GPT-5.6 Sol) still roughly doubles GLM-5.3's exploitation throughput. An open model getting scarier with every release cycle is now a trend, not an incident.
Moonshot AI is racing toward a Hong Kong IPO as the next big Chinese LLM lab to test public markets — with reports of a filing before September 30 piling up against official denials. Market reports in early August said Moonshot planned to file as soon as this month and raise about $3 billion; the company called that untrue, and repeated the denial again this week. Reports say Moonshot is restructuring to bring in state capital and seeking Beijing's approval, while a roughly $50 billion valuation — propelled by Kimi K3's strong Artificial Analysis showing — gives it a premium window. The pressure is timing: DeepSeek's V4 Pro, released August 12, costs under a tenth of Kimi K3 per task under current pricing, undercutting the "frontier plus cheap" story Moonshot would take to market, while Zhipu and MiniMax already listed and used follow-on placements to fund compute. The window may not wait for a perfect narrative.
What to watch: the GLM-5.3 open-weights drop in two weeks — and whether the exploitation curve keeps climbing with it.
If open-weights models keep getting sharper at exploitation with every release, should labs hold weights back when safety hardening lags? Tell us in the comments.
Sources: Z.ai blog · Unite.AI · Bloomberg · Sina Finance (新浪财经) · 21财经 · 36Kr