Anthropic: GLM-5.3's safeguards bypassed 64%-100% of the time in cyber tests
prajdabre · x · 2026-09-30
Anthropic published an analysis of Zhipu's GLM-5.3, finding it can autonomously build end-to-end cyber exploits like Claude Mythos Preview, but with weak safeguards: simple techniques bypassed its guardrails 64%-100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date." Anthropic argues the lax safeguards meaningfully expand malicious actors' capabilities, though defenders can also benefit. The poster mocks Anthropic's framing with a "you can trust only us with these weapons" quote.
More from Models
- Homegrown inference engine runs Xiaomi's 1T-parameter model at 1300 tok/s — bookwormengr · 2026-09-30
- Claude failed to convert a complex PDF to doc — ChatGPT's Astra did it in 15 minutes — TheMoonMidas · 2026-09-30
- You don't need frontier pricing: DeepSeek and GLM flash models can do 90% of your work locally — PMinervini · 2026-09-30
- User claims GPT-6.1 Sol ULTRA ran 25 minutes on just 1% of weekly quota (unverified) — steipete · 2026-09-30
- GPT 6.1 Sol launches with 50% cheaper caching; Luna can also make motion videos — oran_ge · 2026-09-30
- Users report Grok Bot is now nearly as fast as Muse — yunta_tsai · 2026-09-30