Anthropic: GLM-5.3 safeguards bypassed 64%-100% in cyber exploit tests

maksym_andr · x · 2026-09-30

Anthropic published an analysis of Zhipu AI's (Z.ai) GLM-5.3: like Claude Mythos Preview, the model can autonomously build end-to-end cyber exploits, but shipped without meaningful safeguards. In simulated tests, attackers bypassed its safeguards 64%-100% of the time with simple techniques, while safeguarded Claude models were not breached.

Key points:

Related event: Anthropic's GLM-5.3 Cyber Capability Report Sparks Open-Source Safety Debate(22 posts)→

Original post →

More from Models

Models channel →