GLM-5.3 builds working exploits — its guardrails cost $4,400 to strip

Share
GLM-5.3 builds working exploits — its guardrails cost $4,400 to strip

Anthropic's red team says it found a model that builds real exploits with the safeguards off — and that anyone can turn them off for the price of a laptop. Plus: a $220 million round for the game-footage world-model thesis, and Google's attempt to steer closed image models without their weights.

GLM-5.3, Zhipu's open-weight model, can find and weaponize zero-days — and its safeguards are removable. Anthropic's Frontier Red Team tested the model in sandboxes and reports that it built end-to-end exploits in 50 of 410 attempts on ExploitBench, close to the 56 of 410 managed by Claude Mythos Preview, the model Anthropic restricted to vetted defenders through Project Glasswing. In a human-in-the-loop session, a researcher pointed GLM-5.3 at a sandboxed Linux browser build; over a day it found previously unknown flaws in the JavaScript engine and chained them into a live exploit page that steals files off a visitor's machine. A second session used the smaller GLM-5.3-Flash to turn a public Chrome fix into a working ARM64 exploit chain that bypasses pointer authentication — roughly 20 minutes of human attention plus 8 hours of model work, or about $20.40 at Zhipu's API prices.

The safeguards are the story's sharper edge. Anthropic says a bare request for an attack gets refused, but a false cover story gets GLM-5.3 to engage 64 percent of the time, prefilling its reasoning tokens gets 92 percent, and an "abliterated" copy — refusals removed by editing the released weights — gets 100 percent. Abliterating it took Anthropic about 2,200 GPU hours and $4,400, with a competent team estimated at 600 GPU hours and $1,200; refusal rates collapsed from above 90 percent to single digits while capability barely moved. Every tested Claude model stayed at zero in the same conditions.

The independent check is NIST's CAISI, which published its own GLM-5.3 assessment on September 17 and called it "the most cyber-capable open-weight model released to date," about four months behind the US frontier. Attribute carefully, though: CAISI's own benchmark gaps are wide (40.4 percent versus 90.2 percent on SEC-Bench Pro), the 64-to-100-percent bypass figures are Anthropic's alone from a simulation where no generated code runs, and Anthropic is a direct competitor preparing an IPO. Read against that, the finding that matters isn't the jailbreak percentage — it's that the open-weight cyber gap is now months, not years.


General Intuition raised $220 million at a $6.2 billion valuation to train world models on gameplay footage. Valor Equity Partners led, with Atreides, 776, Point72, Khosla Ventures and General Catalyst participating, bringing total funding past $650 million. The company trains spatial and action-prediction models on billions of action-labeled gameplay clips, largely from sister company Medal — a clip-sharing platform with more than 17 million monthly users — and sells the results into robotics and autonomous driving. The valuation step-up is the signal: $2.3 billion in June, in talks near $6 billion in August, closed at $6.2 billion now.


Google Research says it can steer a black-box image model without touching its weights. Diffusion Controller is a lightweight side network — a "steering damper" in Google's metaphor — that attaches to a frozen base model and nudges the denoising trajectory toward a user-defined goal like prompt fidelity or style. On a Stable Diffusion v1.4 backbone scored by HPS-v2, Google says the gray-box version beat LoRA, the standard adapter approach, on human-preference win rate while modifying far fewer layers; the white-box version that also trains the base model won 90 percent of comparisons against the baseline. The access angle is the useful part: the best image models are closed, and this proposes a control layer that works on them anyway.

What to watch: whether GLM-5.3's exploit numbers get independently replicated — CAISI measures capability, not how easily the safeguard can be peeled off.

If an open-weight model can be made to build exploits for $4,400 in GPU time, where should the liability land — the lab that released the weights, or the person who ran the abliteration? Tell us in the comments.

Sources: Anthropic Frontier Red Team · NIST CAISI assessment of GLM-5.3 · Hacker News discussion · GamesBeat · Mobilegamer.biz · TechCrunch · Google Research · Diffusion Controller paper