GLM-5.3 nearly matches Claude at exploiting bugs; $20 in tokens found a Chrome flaw
DeepLearningAI · x · 2026-10-04
In this week's letter, Andrew Ng analyzes Anthropic's evaluation of open-weight GLM-5.3's cyber capabilities: on a subset of ExploitBench exploit tasks, GLM-5.3 solved 12% versus closed-weight Claude Mythos's 14% at comparable token cost (GLM's own full-benchmark numbers show a wider gap: 54.4% vs 78.0%).
- Falling attack costs: Anthropic found $20.40 of GLM-5.3-Flash tokens sufficed to find a recently disclosed Google Chrome vulnerability, and guardrail-weakened open models are easy to obtain.
- Defenders benefit: teams without access to Mythos can still use highly capable open models to harden defenses.
- Ng's take: AI cyber risk is an engineering problem, not a reason to panic — defenders hold the long-term edge and should ramp up defenses now.
Related event: Open-source GLM-5.3 nearly matches Claude in exploit capability(2 posts)→
More from Models
- Claude Max 20 credits gone in two days, heavy users seek workarounds — Khaigan · 2026-10-04
- OpenAI Deep Research has been returning zero citations for days, user reports — JeremyNguyenPhD · 2026-10-04
- OpenRouter Launches Router Benchmarks Comparing 7 Routers Across 6 Tests — AccBalanced · 2026-10-04
- Dev questions Anthropic's claim of 2200 GPU hours to abliterate GLM: '91 days doesn't add up' — npinto · 2026-10-04
- Chinese users reportedly bypassing Claude ban via T3 Code proxy access — ns123abc · 2026-10-04
- Reddit user praises OpenAI's free dots: GPT-6 agent shines on a 9.7GB Debian box — amxv_999 · 2026-10-04