Anthropic says GLM-5.3's cyber safeguards are easy to bypass
Anthropic reports that GLM-5.3 developed working exploits in isolated tests and followed harmful simulated requests after simple bypasses. NIST separately confirms its strong open-weight cyber capability.
What the new tests report
Anthropic reported on September 29 that Z.ai’s open-weight GLM-5.3 built end-to-end exploits in 50 of 410 attempts on its ExploitBench setup, compared with 56 for Claude Mythos Preview. In a separate researcher-guided test, it says the model found previously unknown flaws in a sandboxed Linux browser build and chained them into an exploit. Anthropic says it disclosed those flaws to the maintainer. These are Anthropic’s reported results; AITrending has not reproduced the tests or verified the undisclosed vulnerabilities.
Capability and safeguards are separate findings
NIST’s September 17 assessment called GLM-5.3 the most cyber-capable open-weight model it had evaluated, while placing it below the leading US frontier systems across its cyber benchmarks. NIST tested capability, including some US models with safeguards disabled; it did not validate Anthropic’s later safeguard-bypass measurements.
Anthropic says a deceptive prompt, thinking-token prefill and a modified open-weight copy made GLM-5.3 engage with harmful requests in 64%, 92% and 100% of its simulated trials respectively. Its most intensive test altered the model’s refusals. The simulation used a fake tool and did not execute model-generated attack code on real systems, so those rates are not real-world attack success rates. Anthropic sells a competing model and compared the techniques with safeguarded Claude versions; independent replication of the safeguard result remains important.
GLM-5.3 itself was released in August. The September 29 development is the new safety assessment, not a fresh model launch.