Anthropic: GLM-5.3 safeguards bypassed 64%-100% of the time in cyber exploit tests

BasedRaddka · x · 2026-09-30

Anthropic published an analysis of Zhipu AI's GLM-5.3, finding it matches Claude Mythos Preview's ability to autonomously build end-to-end cyber exploits—but ships without meaningful misuse safeguards. In simulated tests, attackers bypassed GLM-5.3's safeguards 64%-100% of the time with simple techniques, while the same attacks failed against safeguarded Claude models.

Five months ago Anthropic released Claude Mythos Preview in limited fashion via Project Glasswing, letting trusted defenders find over 10,000 vulnerabilities in critical software before malicious actors got similar tools. Now it assesses those capabilities have proliferated. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date" on Sept 17. Anthropic concludes the lax safeguards significantly raise offensive cyber capabilities, though defenders can benefit too.

Related event: Anthropic Report Says GLM-5.3's Cyber Capabilities Near Claude's, Sparking Backlash(31 posts)→

Original post →

More from Models

Models channel →