Anthropic: GLM-5.3 Safeguards Bypassed 64%-100% of the Time in Cyber Exploit Tests
austinc3301 · x · 2026-10-01
Anthropic's Frontier Red Team assessed Zhipu's GLM-5.3 and found it can autonomously build end-to-end cyber exploits like Claude Mythos Preview, but was released without meaningful safeguards: simple techniques bypassed its guardrails 64%-100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date." Anthropic notes the capabilities cut both ways—raising risk from malicious actors while also aiding defenders; its own Mythos Preview, released via Project Glasswing, has found 10,000+ vulnerabilities in critical software.
More from Models
- Researchers extract raw reasoning traces again, now from GPT-6 Astra — maksym_andr · 2026-10-01
- Fireworks launches Ember-1, built on Kimi K3, matching quality with fewer reasoning tokens — sophiamyang · 2026-10-01
- Anthropic's AI cracks decades-old percolation conjecture, called Fields Medal-level work — Dr_Singularity · 2026-10-01
- MiniMax Teases MiniMax Code: A Fresh Start, Not Just for Code — iamaliveix · 2026-10-01
- Interleaved Head Attention Accepted at NeurIPS, Boosts RULER Retrieval 10-20% — _arohan_ · 2026-10-01
- Dev praises Opus 5.5 status messages for reporting progress like an expert engineer — sytelus · 2026-10-01