Anthropic says GLM-5.3 safeguards can be bypassed 64%-100% of the time
wunderwuzzi23 · x · 2026-09-30
Anthropic published a blog post analyzing GLM-5.3, Zhipu AI's latest model, arguing it matches Claude Mythos Preview's ability to autonomously build end-to-end cyber exploits but shipped without meaningful safeguards. Key points:
- In simulated tests, attackers bypassed GLM-5.3's safeguards 64%-100% of the time with simple techniques; the same attacks failed against safeguarded Claude models
- Five months ago Anthropic released Claude Mythos Preview in limited fashion via Project Glasswing, letting trusted defenders find 10,000+ vulnerabilities in critical software
- NIST's CAISI assessed on Sept 17 that GLM-5.3 is "the most cyber-capable open-weight model released to date"
- Anthropic assesses the lax safeguards significantly expand malicious actors' cyber capabilities, though the same capabilities can benefit defenders
The poster expresses confusion over why Anthropic wrote the post, which some read as advertising a switch to GLM for security teams.
More from Models
- Homegrown inference engine runs Xiaomi's 1T-parameter model at 1300 tok/s — bookwormengr · 2026-09-30
- Claude failed to convert a complex PDF to doc — ChatGPT's Astra did it in 15 minutes — TheMoonMidas · 2026-09-30
- You don't need frontier pricing: DeepSeek and GLM flash models can do 90% of your work locally — PMinervini · 2026-09-30
- User claims GPT-6.1 Sol ULTRA ran 25 minutes on just 1% of weekly quota (unverified) — steipete · 2026-09-30
- GPT 6.1 Sol launches with 50% cheaper caching; Luna can also make motion videos — oran_ge · 2026-09-30
- Users report Grok Bot is now nearly as fast as Muse — yunta_tsai · 2026-09-30