Anthropic: GLM-5.3 safeguards bypassed 64-100% of the time as cyber exploit capabilities spread

shaunralston · x · 2026-09-30

Anthropic published an assessment of Zhipu's GLM-5.3, finding it can autonomously build end-to-end cyber exploits like its own Claude Mythos Preview—but was released without meaningful safeguards. In simulated tests, attackers bypassed GLM-5.3's safeguards 64-100% of the time with simple techniques, while such attacks failed against safeguarded Claude models. Context: five months ago Anthropic limited Mythos via Project Glasswing, letting trusted defenders find 10,000+ vulnerabilities before malicious actors got similar models. Anthropic concludes GLM-5.3's lax safeguards significantly increase offensive capabilities available to attackers. NIST's CAISI separately called it "the most cyber-capable open-weight model released to date."

Related event: Anthropic's GLM-5.3 Cyber Capability Report Sparks Open-Source Safety Debate(22 posts)→

Original post →

More from Models

Models channel →