Anthropic red team: GLM-5.3 pulls off full control flow hijacks in 4% of binary exploitation trials
Simon Willison · rss · 2026-09-30
Anthropic's Frontier Red Team evaluated models on 100 random tasks from its internal Binary Exploitation benchmark: GLM-5.3 achieved full control flow hijacks in 4% of trials and Claude Mythos Preview in 6%, while earlier models like Claude Opus 4.6 and GLM-5.2 succeeded in none. The report argues a meaningful threshold in advanced cyber capability has been crossed.
More from Safety
- Meta's Muse AI accused of uploading Apple Messages to the cloud even when users opt out — SumitGup · 2026-09-30
- Math advisory group issues responsible-release guidelines for AI-generated math; Gowers amplifies, critics push back — RexDouglass · 2026-09-30
- AI Lawsuit Risks Becoming a Court-Ordered Kill Switch While Chinese Rivals Race Ahead — castrotech · 2026-09-30
- Apollo Research shifts to embedded evaluators with employee-equivalent access — MariusHobbhahn · 2026-09-30
- Trump and top AI executives sign 'morally binding' voluntary AI controls — sourdub · 2026-09-30
- Texas police audit finds 'egregious misuse' of Flock cameras, dispatcher tracked her child 42 times — Polymarket · 2026-09-30