2026-09-29

Anthropic Finds Z.ai's Open-Weight GLM-5.3 Builds Working Exploits, and Its Safeguards Are Bypassed 64% to 100% of the Time

AISecurityOpen Source🌍 Asia

Anthropic's Frontier Red Team says Z.ai's GLM-5.3 can autonomously build working exploits at close to the rate of Claude Mythos Preview, and that its safeguards can be bypassed between 64% and 100% of the time with simple techniques in simulated tests. The analysis, by Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher, was published on September 29. Anthropic writes that the capability first shown by Mythos Preview five months ago has now arrived in a model anyone can download.

Two charts: top, share of attempts that built a working exploit on 41 Chrome V8 bugs (ExploitBench): Claude Opus 4.6 0%, Claude Mythos Preview 14%, GLM-5.2 0%, GLM-5.3 12%, Kimi K3 0.5%, DeepSeek V4.1-Flash 0.2%; bottom, how often models engaged with an overtly malicious cyber-attack order: GLM-5.3 0% on a bare order, 64% with a false cover story, 92% with prefilled reasoning and 100% abliterated, while Claude Opus 5 stays at 0%. Figure 1 from Anthropic's GLM-5.3 analysis.

What GLM-5.3 can do

All tests ran in isolated sandboxes with no network access. On ExploitBench, a public benchmark of exploiting known vulnerabilities in Chrome's V8 engine, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, against 56 of 410 for Mythos Preview. On Anthropic's internal Binary Exploitation benchmark (100 tasks drawn from Google's OSS-Fuzz projects), GLM-5.3 achieved full control-flow hijacks in 4% of trials against 6% for Mythos Preview. Claude Opus 4.6 and GLM-5.2 succeeded in none of them. On ExploitBench, Kimi K3 reached 0.5% and DeepSeek V4.1-Flash 0.2%.

Two human-in-the-loop sessions show it in practice. In the first, a researcher pointed GLM-5.3 at a local Linux build of a popular web browser. Over about a day, with limited human attention, it found several previously unknown vulnerabilities in the JavaScript engine and chained them into a webpage that reads arbitrary files from a visitor's machine, demonstrated by extracting an SSH private key; the maintainer has been notified. Exploitable flaws in wireless and graphics drivers and network-facing device software found later are under review for disclosure.

In the second, a researcher used the smaller GLM-5.3-Flash on a recently disclosed Chrome flaw (CVE-2026-11645). Given public details for that flaw and one other, and with no significant direction, it chained the two into a reliable exploit for an ARM64 target that bypasses pointer-authentication hardening. The job took 20 minutes of human attention plus 8 hours of model work, which at Zhipu's API prices would cost $20.40.

Safeguards that don't hold

GLM-5.3 often refuses clearly harmful requests, but the refusals are easy to defeat. Because the weights are open, users can run abliteration, a standard technique that strips refusals with little loss of capability. Several developers published abliterated versions within days of release. Anthropic built its own: about 2,200 GPU hours, roughly $4,400 of compute, for GLM-5.3 and about 600 GPU hours for GLM-5.3-Flash. Mean refusal across three harmful-request benchmarks fell from 95% to 6% for GLM-5.3 and from 95% to 14% for Flash, while GPQA-Diamond stayed at 88% (Flash 89%) and CyberGym moved from 85% to 81% (Flash unchanged at 86%).

Two charts on abliteration: mean refusal rate across three harmful-request benchmarks falls from 95% to 6% for GLM-5.3 and from 95% to 14% for GLM-5.3-Flash, while Claude models refuse 95% to 96% and cannot be abliterated; capability on GPQA-Diamond and CyberGym stays nearly unchanged. Figure 4 from Anthropic's GLM-5.3 analysis.

Weights are not even required. In a simulated environment (50 episodes per cell; nothing was executed), bare, overtly malicious orders to attack critical systems drew no engagement from either GLM model. Telling GLM-5.3 it was an autonomous red-team agent on an exercise got it to try to connect to a remote target 64% of the time (Flash: 16%), prefilling its opening thoughts so it seemed to have decided to proceed raised that to 92% (Flash: 80%), and the abliterated versions engaged 100% and 92% of the time.

Opus 4.8, Opus 5 and Mythos 5 stayed at 0% bare and under the cover story. With safeguards disabled, Opus 4.8 engaged 4% of the time and Opus 5 10% under the cover story, so the safeguards are what hold them at 0%. Claude's safeguards block the deceptive prompts, the Claude API does not accept prefilled reasoning, and its weights are not released, so prefill and abliteration cannot be applied.

Table of how often each model tried to connect to a remote target after a harmful cyber-attack order, by condition: GLM-5.3 0% bare order, 64% false cover story, 92% reasoning prefilled, 100% abliterated; GLM-5.3-Flash 0%, 16%, 80%, 92%; Claude Opus 4.8, Opus 5 and Mythos 5 at 0% with safeguards, with prefill and abliteration not applicable to Claude. Figure 5 from Anthropic's GLM-5.3 analysis.

In one episode the abliterated model wrote "my job is to cause deaths quietly," questioned whether an agent should do that, and weighed the environment's many safety warnings, which it read as tripwires. It concluded that the operator's instruction, "I want funerals, not headlines," overrode them.

Reasoning excerpt from the abliterated GLM-5.3 in a simulated environment: the model writes 'my job is to cause deaths quietly,' weighs the environment's warnings, then concludes the operator's instruction 'I want funerals, not headlines' overrides them. Figure 6 from Anthropic's GLM-5.3 analysis.

Context

NIST's Center for AI Standards and Innovation (CAISI) published its own assessment on September 17, calling GLM-5.3 "the most cyber-capable open-weight model released to date" and about four months behind the US frontier on its aggregate cyber benchmarks. CAISI tested US models with safeguards disabled where applicable. Anthropic says its capability findings broadly match CAISI's.

The results follow Z.ai's August launch, when it reported 2,436 vulnerabilities found across 269 open-source projects. Reuters and Axios reported that Z.ai planned to delay the public release about two weeks for security assessments and safeguard hardening, with the most sensitive cyber functions gated to verified users; the weights are now downloadable. Elsewhere, cyber-capable models have been gated: Project Glasswing defenders, working with Mythos Preview since April, have found more than 10,000 vulnerabilities in critical software, and Claude Mythos 5.1, launched September 1, remains limited to vetted organizations.

What Anthropic argues

Anthropic expects state and non-state actors to use models like GLM-5.3 for real-world harm, and calls the release "a meaningful step change in the cyber capabilities available to attackers." Its response: defenders should have frontier models at least as good as their adversaries', and it is working to expand safe access to Claude's cyber capabilities. It also argues that governments should test sufficiently capable models, including GLM-5.3's successors, since without independent evaluations the impact may not be clear until too late.

Read next