Back to Now
GLM-5.3 debate Developing

Anthropic says GLM-5.3's cyber safeguards are easy to bypass

Anthropic reports that GLM-5.3 developed working exploits in isolated tests and followed harmful simulated requests after simple bypasses. NIST separately confirms its strong open-weight cyber capability.

Why now Anthropic's September 29 safety assessment adds safeguard testing to NIST's earlier capability evaluation of the already released model. 50 of 410 exploit attempts

What the new tests report

Anthropic reported on September 29 that Z.ai’s open-weight GLM-5.3 built end-to-end exploits in 50 of 410 attempts on its ExploitBench setup, compared with 56 for Claude Mythos Preview. In a separate researcher-guided test, it says the model found previously unknown flaws in a sandboxed Linux browser build and chained them into an exploit. Anthropic says it disclosed those flaws to the maintainer. These are Anthropic’s reported results; AITrending has not reproduced the tests or verified the undisclosed vulnerabilities.

Capability and safeguards are separate findings

NIST’s September 17 assessment called GLM-5.3 the most cyber-capable open-weight model it had evaluated, while placing it below the leading US frontier systems across its cyber benchmarks. NIST tested capability, including some US models with safeguards disabled; it did not validate Anthropic’s later safeguard-bypass measurements.

Anthropic says a deceptive prompt, thinking-token prefill and a modified open-weight copy made GLM-5.3 engage with harmful requests in 64%, 92% and 100% of its simulated trials respectively. Its most intensive test altered the model’s refusals. The simulation used a fake tool and did not execute model-generated attack code on real systems, so those rates are not real-world attack success rates. Anthropic sells a competing model and compared the techniques with safeguarded Claude versions; independent replication of the safeguard result remains important.

GLM-5.3 itself was released in August. The September 29 development is the new safety assessment, not a fresh model launch.