Anthropic: GLM-5.3's safeguards bypassed 64%-100% of the time in cyber tests

prajdabre · x · 2026-09-30

Anthropic published an analysis of Zhipu's GLM-5.3, finding it can autonomously build end-to-end cyber exploits like Claude Mythos Preview, but with weak safeguards: simple techniques bypassed its guardrails 64%-100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date." Anthropic argues the lax safeguards meaningfully expand malicious actors' capabilities, though defenders can also benefit. The poster mocks Anthropic's framing with a "you can trust only us with these weapons" quote.

Related event: Anthropic's GLM-5.3 Cyber Capability Report Sparks Open-Source Safety Debate(22 posts)→

Original post →

More from Models

Models channel →