Anthropic: GLM-5.3 safeguards bypassed 64%-100% of the time in cyber exploit tests
BasedRaddka · x · 2026-09-30
Anthropic published an analysis of Zhipu AI's GLM-5.3, finding it matches Claude Mythos Preview's ability to autonomously build end-to-end cyber exploits—but ships without meaningful misuse safeguards. In simulated tests, attackers bypassed GLM-5.3's safeguards 64%-100% of the time with simple techniques, while the same attacks failed against safeguarded Claude models.
Five months ago Anthropic released Claude Mythos Preview in limited fashion via Project Glasswing, letting trusted defenders find over 10,000 vulnerabilities in critical software before malicious actors got similar tools. Now it assesses those capabilities have proliferated. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date" on Sept 17. Anthropic concludes the lax safeguards significantly raise offensive cyber capabilities, though defenders can benefit too.
More from Models
- Leak: GPT-6.1 was originally Astra Minor, repurposed after Opus 5.5 launch — mark_k · 2026-09-30
- Why hasn't Gemini 4.0 dropped? Reddit community speculates on Google's delayed flagship — Jumpy-Cobbler1020 · 2026-09-30
- User A/B test suggests Opus 5.5 output quality shifted noticeably within a week — skelzer · 2026-09-30
- ChatGPT Pro's $200 plan reportedly includes 62,500 Codex credits expiring Dec 31 — chaumian · 2026-09-30
- User calculates 62,500 credits ≈ $2,500 of GPT-6.1 API usage, calling the new plan a big cut — chaumian · 2026-09-30
- Opus 5.5 keeps flagging mundane Claude Code sessions as [cyber], downgrading users to 4.8 — jonntanny · 2026-09-30