Anthropic: GLM-5.3 safeguards bypassed 64%-100% in cyber exploit tests
maksym_andr · x · 2026-09-30
Anthropic published an analysis of Zhipu AI's (Z.ai) GLM-5.3: like Claude Mythos Preview, the model can autonomously build end-to-end cyber exploits, but shipped without meaningful safeguards. In simulated tests, attackers bypassed its safeguards 64%-100% of the time with simple techniques, while safeguarded Claude models were not breached.
Key points:
- Five months ago Anthropic released Mythos Preview in limited fashion via Project Glasswing, helping trusted defenders find 10,000+ vulnerabilities in critical software.
- On Sept 17, NIST's CAISI independently assessed GLM-5.3 as "the most cyber-capable open-weight model released to date."
- Anthropic assesses GLM-5.3's lax safeguards significantly expand offensive capabilities for malicious actors, though defenders can also benefit.
More from Models
- Homegrown inference engine runs Xiaomi's 1T-parameter model at 1300 tok/s — bookwormengr · 2026-09-30
- Claude failed to convert a complex PDF to doc — ChatGPT's Astra did it in 15 minutes — TheMoonMidas · 2026-09-30
- You don't need frontier pricing: DeepSeek and GLM flash models can do 90% of your work locally — PMinervini · 2026-09-30
- User claims GPT-6.1 Sol ULTRA ran 25 minutes on just 1% of weekly quota (unverified) — steipete · 2026-09-30
- GPT 6.1 Sol launches with 50% cheaper caching; Luna can also make motion videos — oran_ge · 2026-09-30
- Users report Grok Bot is now nearly as fast as Muse — yunta_tsai · 2026-09-30