Anthropic: GLM-5.3 Safeguards Bypassed 64%-100% of the Time in Cyber Exploit Tests

austinc3301 · x · 2026-10-01

Anthropic's Frontier Red Team assessed Zhipu's GLM-5.3 and found it can autonomously build end-to-end cyber exploits like Claude Mythos Preview, but was released without meaningful safeguards: simple techniques bypassed its guardrails 64%-100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date." Anthropic notes the capabilities cut both ways—raising risk from malicious actors while also aiding defenders; its own Mythos Preview, released via Project Glasswing, has found 10,000+ vulnerabilities in critical software.

Related event: Anthropic Report Warns GLM-5.3 Near-Frontier Cyber Capabilities, Sparking Debate(37 posts)→

Original post →

More from Models

Models channel →