Anthropic: GLM-5.3 safeguards bypassed 64-100% of the time as cyber exploit capabilities spread
shaunralston · x · 2026-09-30
Anthropic published an assessment of Zhipu's GLM-5.3, finding it can autonomously build end-to-end cyber exploits like its own Claude Mythos Preview—but was released without meaningful safeguards. In simulated tests, attackers bypassed GLM-5.3's safeguards 64-100% of the time with simple techniques, while such attacks failed against safeguarded Claude models. Context: five months ago Anthropic limited Mythos via Project Glasswing, letting trusted defenders find 10,000+ vulnerabilities before malicious actors got similar models. Anthropic concludes GLM-5.3's lax safeguards significantly increase offensive capabilities available to attackers. NIST's CAISI separately called it "the most cyber-capable open-weight model released to date."
More from Models
- Homegrown inference engine runs Xiaomi's 1T-parameter model at 1300 tok/s — bookwormengr · 2026-09-30
- Claude failed to convert a complex PDF to doc — ChatGPT's Astra did it in 15 minutes — TheMoonMidas · 2026-09-30
- You don't need frontier pricing: DeepSeek and GLM flash models can do 90% of your work locally — PMinervini · 2026-09-30
- User claims GPT-6.1 Sol ULTRA ran 25 minutes on just 1% of weekly quota (unverified) — steipete · 2026-09-30
- GPT 6.1 Sol launches with 50% cheaper caching; Luna can also make motion videos — oran_ge · 2026-09-30
- Users report Grok Bot is now nearly as fast as Muse — yunta_tsai · 2026-09-30