Anthropic red team: GLM-5.3 safeguards bypassed 64%-100% in simulated cyber tests
dl_weekly · x · 2026-10-09
Anthropic's Frontier Red Team analyzed GLM-5.3, Zhipu AI's (Z.ai) latest model, finding it matches Claude Mythos Preview's ability to autonomously build end-to-end cyber exploits—but shipped without meaningful safeguards.
- Simple techniques bypassed GLM-5.3's safeguards 64%–100% of the time in simulated tests; the same attacks failed against safeguarded Claude models.
- Five months ago Anthropic released Mythos Preview in limited fashion via Project Glasswing, letting trusted defenders find 10,000+ vulnerabilities in critical software; it warned such capabilities would proliferate.
- NIST's CAISI independently assessed GLM-5.3 on Sept. 17 as "the most cyber-capable open-weight model released to date."
- Anthropic assesses the lax safeguards significantly expand malicious actors' cyber capabilities, though the same abilities can also aid defenders.
More from Models
- LightOnOCR-3 released: 0.8B/1B/4B open models for OCR, layout, charts under Apache 2.0 — antoine_chaffin · 2026-10-09
- Perplexity Decider tops DecisionBench with 93.9% accuracy, 534ms latency, and lowest cost — AravSrinivas · 2026-10-09
- Claude 3 Opus confirmed working within Max plan monthly API credits — repligate · 2026-10-09
- Leak: Grok Voice Mode coming to X — talk to Grok out loud in the app — nima_owji · 2026-10-09
- Subscription Claude models deliver far fewer thinking tokens, measured five ways — _AustinCalvert_ · 2026-10-09
- GPT-6 in ChatGPT is the best AI writer yet — barely needs editing, says user — VraserX · 2026-10-09