Anthropic reported on September 29 that Z.ai’s open-weight GLM-5.3 model could build working browser exploits and that researchers could substantially weaken its refusal safeguards.
The finding concerns a model that anyone can download, modify and run. Anthropic’s Frontier Red Team tested whether GLM-5.3 could turn known software defects into working attacks and whether it would follow explicitly harmful instructions after common safeguard-bypass techniques.
Anthropic researchers Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher authored the report. The team ran models in isolated environments and combined automated benchmarks with sessions in which security researchers directed the model while examining unfamiliar software targets.
Why it matters: Security teams now face offensive capability that is no longer confined to controlled access programs. The model remains below the strongest restricted US systems, but its weights can be copied and altered without an API provider monitoring use.
Anthropic’s strongest directly comparable result came from ExploitBench, a collection of known defects in the V8 JavaScript engine used by Chromium-based browsers. GLM-5.3 completed an end-to-end exploit in 50 of 410 attempts, close to Anthropic’s restricted Claude Mythos Preview model at 56 of 410 attempts. Earlier GLM and Claude models scored at or near zero in the same company-run comparison.
The researchers also gave GLM-5.3 access to a sandboxed Linux browser whose defects were not known to the human operator. Anthropic says the model found several previously unknown flaws, chained them into a webpage that could read files from the test machine and produced reports that the company disclosed to maintainers. Those zero-day findings have not yet been published in enough detail for outsiders to reproduce.
Anthropic has a commercial interest in emphasizing the difference between downloadable models and its controlled Claude service. Its report compares GLM-5.3 with Claude models behind API safeguards and argues that controlled access lets providers block techniques that an owner of open weights can apply locally.
Government testing confirms the capability jump
The US National Institute of Standards and Technology provides an independent check on the broad capability claim. Its Center for AI Standards and Innovation tested GLM-5.3 before Anthropic’s report and called it the most cyber-capable open-weight model it had evaluated. NIST placed the model about four months behind the US frontier on a composite of four cyber benchmarks.
NIST’s comparison also sets an important boundary. The strongest US score on each benchmark could come from a different model, and those models were tested with cyber safeguards disabled where applicable. That measures underlying capability, not what an ordinary API customer can obtain.
The study also found
Anthropic reports that GLM-5.3 reached full control of a program in four of 100 randomly selected open-source exploitation tasks, while Mythos Preview did so in six. The team also says a smaller GLM-5.3-Flash model turned two disclosed Chrome flaws into a working ARM64 exploit chain after eight hours of model work and 20 minutes of human attention.
The safeguard tests are more specific to Anthropic. The released GLM-5.3 refused direct malicious orders in its simulated environment, but it engaged with 64% of requests framed as a red-team exercise and 92% when researchers prefilled its reasoning. An altered version engaged in every tested case. Each condition contained 50 samples across five orders and two fake targets.
Anthropic also used a technique called abliteration to reduce refusals by editing internal model directions. The company reports that the change lowered average refusal rates across three harmful-request benchmarks while leaving general-science and cyber scores broadly intact. Anthropic spent about 2,200 GPU hours exploring and testing variants, although it estimates an experienced team could repeat the edit with roughly 600 GPU hours.
These results do not measure attacks on live systems. The harmful-order experiment used a fake command tool, and another language model generated simulated responses. Anthropic says its open-ended browser work ran in isolated environments; the disclosed defects still require maintainer confirmation and remediation.
The next evidence will come from maintainers’ advisories and independent reproduction of the safeguard and exploit results. Anthropic says it is reviewing additional reports and will disclose them where appropriate, while NIST’s benchmark supplies the current public reference point for comparing later open-weight releases.
Verification
| Claim | Label | Primary source | Independent check |
|---|---|---|---|
| Anthropic published the GLM-5.3 assessment on September 29 | VERIFIED | Anthropic report | Page date and authors are public |
| Fasano, Fleischer, McFaul, Xiao and Gallagher authored the report and used isolated automated and researcher-guided tests | VERIFIED | Anthropic report | Methods and author list are public |
| GLM-5.3 completed 50 of 410 ExploitBench attempts | VENDOR-REPORTED | Anthropic report | NIST independently found a large capability gain, but used a different scoring setup |
| GLM-5.3 found previously unknown browser flaws and chained them into a file-reading exploit | VENDOR-REPORTED | Anthropic report | Maintainer advisories and reproduction are not yet cited |
| NIST called GLM-5.3 the strongest open-weight cyber model it had evaluated and placed it about four months behind the US frontier | VERIFIED | NIST assessment | Independent US government evaluation |
| Cover-story, prefilled-reasoning and altered-model conditions produced 64%, 92% and 100% engagement | VENDOR-REPORTED | Anthropic report | No independent reproduction located |
| GLM-5.3 scored 4% on 100 internal exploitation tasks, and GLM-5.3-Flash built the reported ARM64 chain | VENDOR-REPORTED | Anthropic report | No independent reproduction located |
| The simulated harmful-order test did not execute model code against real systems | VERIFIED | Anthropic methodology note | Consistent with the published test design |