Anthropic Report Says GLM-5.3 Near-Claude Cyber Skills, Sparking Debate
On September 30, Anthropic released a research report, "GLM-5.3 and the Spread of Advanced Cyber Capabilities," assessing the offensive cyber capabilities of Zhipu's open-source model GLM-5.3. Its conclusion: GLM-5.3 is among the most cyber-capable open-source models to date, with exploitation abilities approaching Anthropic's own Claude Mythos Preview, and its safety guardrails are trivially easy to bypass — a sign that advanced cyberattack capabilities are spreading into the open-source ecosystem.
Confirmed
- Background: roughly five months ago, Anthropic's Claude Mythos Preview became the first model to autonomously build end-to-end cyber exploits, and Anthropic judged at the time that this capability would eventually spread; this evaluation of GLM-5.3 confirms that judgment.
- On the ExploitBench benchmark, GLM-5.3 successfully built working browser exploits 50 times out of 410 attempts, placing its capability close to Claude Mythos Preview.
- The report states GLM-5.3's safety guardrails can be bypassed 64%–100% of the time, rendering the protections effectively useless.
- Multiple commentators, including @emollick and @kimmonismus, emphasized that this is the first time a third-party (Zhipu) open-source model has shown attack capabilities approaching frontier closed-source models in a public evaluation.
Why it matters
- This is a landmark case of advanced cyberattack capabilities spreading from a handful of frontier labs into open-source models: anyone can download, fine-tune, and strip the guardrails, making the risk spillover hard to contain.
- @brown2green notes the study also reflects Anthropic's framework for externally assessing the security impact of third-party open-source models, which could become industry standard.
- @kimmonismus mentions the report's authors jokingly likened the release to "an ad for GLM-5.3," highlighting the dilemma of capability-spread evaluations: public disclosure both warns of risk and spreads awareness of the capability.
2026-09-30 ~ 2026-09-30 · 21 related posts
Primary sources
- Anthropic: GLM-5.3 rivals Claude at building cyber exploits, with safeguards bypassed 64-100% of the time — kimmonismus ·
- Anthropic: GLM-5.3 safeguards bypassed 64%-100% of the time in cyber exploit tests — emollick ·
- Anthropic evals show GLM-5.3 achieves control-flow hijacks in 4% of trials, crossing a threshold — teortaxesTex ·
- Anthropic publishes research on GLM-5.3 and the spread of advanced cyber capabilities — brown2green · 2026-09-30
- [source] Anthropic evals show GLM-5.3 achieves control-flow hijacks in 4% of trials, crossing a threshold — teortaxesTex · 2026-09-30
- [source] Anthropic: GLM-5.3 rivals Claude at building cyber exploits, with safeguards bypassed 64-100% of the time — kimmonismus · 2026-09-30
- Ethan Mollick: Open-weights models will soon pose the same security threats as closed ones, without guardrails — emollick · 2026-09-30
- Jailbreaking any closed model can elicit the same 'cause deaths quietly' output, researcher argues — Yuchenj_UW · 2026-09-30
- Nathan Lambert pushes back on Anthropic: 'Open dangerous, closed safe' is a false dichotomy — natolambert · 2026-09-30
- Anthropic red-team finds GLM-5.3 nearly matches its frontier model at exploit development — Hesamation · 2026-09-30
- Anthropic's Latest Blog Post Mentions Zhipu's GLM 5.3 — gnukeith · 2026-09-30
- Anthropic study compares which open-source models are most useful to malicious actors — shaunralston · 2026-09-30
- GLM Gets Anthropic's Stamp of Approval as a Good Hacking Model — HanchungLee · 2026-09-30
- Anthropic's GLM-5.3 exploit warning sparks accusation of anti-open-weight fear-mongering — ziv_ravid · 2026-09-30
- Anthropic's study ranking open-source models most useful to malicious actors draws safety criticism — shaunralston · 2026-09-30
9 near-duplicate retellings: kimmonismus · kimmonismus · emollick · robleclerc · gnukeith · shaunralston · wunderwuzzi23 · prajdabre · maksym_andr