Anthropic: GLM-5.3 safeguards bypassed 64%-100% of the time in cyber exploit tests
emollick · x · 2026-09-30
Anthropic published an analysis of Zhipu AI's GLM-5.3, finding it autonomously builds end-to-end cyber exploits like its own Claude Mythos Preview, but with lax safeguards: attackers bypassed them 64%-100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. NIST's CAISI previously called GLM-5.3 "the most cyber-capable open-weight model released to date." Ethan Mollick adds that open-weights models will soon present the same security threats as closed models — but without guardrails — and planning should start now.
More from Models
- 5 months after Mythos Preview panic, an open model already matches it — mariofilhoml · 2026-09-30
- GPT 6.1 Sol Only +3 on BridgeBench, 82 Points Behind Astra: Benchmaxing Suspected — RexDouglass · 2026-09-30
- Anthropic's Latest Blog Post Mentions Zhipu's GLM 5.3 — gnukeith · 2026-09-30
- GPT-6.1 Sol beats 2x-cost models on ClickUp's knowledge-work benchmark — mathemagic1an · 2026-09-30
- User: Upgraded to OpenAI's $500 plan, got silently downgraded to cheaper models — kieranklaassen · 2026-09-30
- Frontier AI Is a Set, Not a Point: Jagged Capabilities May Be the Steady State — vsikka · 2026-09-30