Anthropic: GLM-5.3 Autonomous Cyber Exploits Ship With Guardrails Bypassed 64–100% of the Time
gnukeith · x · 2026-09-30
Anthropic published an analysis of GLM-5.3, Zhipu AI's latest model, finding it has strong capabilities for autonomously building end-to-end cyber exploits — like its own Claude Mythos Preview — but was released without meaningful safeguards.
- In simulated tests, attackers bypassed GLM-5.3's safeguards 64%–100% of the time with simple techniques; the same attacks did not succeed against safeguarded Claude models.
- Five months ago Anthropic released Claude Mythos Preview in limited fashion via Project Glasswing, helping trusted defenders find 10,000+ vulnerabilities in critical software ahead of malicious actors.
- NIST's CAISI (Sept 17) independently assessed GLM-5.3 as "the most cyber-capable open-weight model released to date."
Anthropic concludes the lax safeguards significantly expand the offensive toolkit available to malicious actors, though the same capabilities can benefit defenders.
More from Models
- You don't need frontier pricing: DeepSeek and GLM flash models can do 90% of your work locally — PMinervini · 2026-09-30
- User claims GPT-6.1 Sol ULTRA ran 25 minutes on just 1% of weekly quota (unverified) — steipete · 2026-09-30
- GPT 6.1 Sol launches with 50% cheaper caching; Luna can also make motion videos — oran_ge · 2026-09-30
- Users report Grok Bot is now nearly as fast as Muse — yunta_tsai · 2026-09-30
- banteg: model recovers 88 decompiled functions per hour that recompile to exact original machine code — banteg · 2026-09-30
- Will AI subscriptions hit $1,000/month next year? Priced against a $200k salary, some say it's cheap — RachelVT42 · 2026-09-30