Anthropic: GLM-5.3 safeguards bypassed 64%-100% of the time in cyber exploit tests

emollick · x · 2026-09-30

Anthropic published an analysis of Zhipu AI's GLM-5.3, finding it autonomously builds end-to-end cyber exploits like its own Claude Mythos Preview, but with lax safeguards: attackers bypassed them 64%-100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. NIST's CAISI previously called GLM-5.3 "the most cyber-capable open-weight model released to date." Ethan Mollick adds that open-weights models will soon present the same security threats as closed models — but without guardrails — and planning should start now.

Related event: Anthropic Report: Open-Source GLM-5.3 Nears Claude's Cyber Attack Capabilities with Easily Bypassed Guardrails(9 posts)→

Original post →

More from Models

Models channel →