Anthropic evals: open-weights GLM-5.3 writes working exploits nearly as well as its restricted Mythos model
lxfater · x · 2026-09-30
Anthropic's safety evaluation of Zhipu's open-weights GLM-5.3 finds cyber-offense capability close to Mythos Preview, an internal high-risk model kept from public release.
- Exploit writing: on the same set of 410 Chrome V8 known-vulnerability tasks per model, GLM-5.3 produced working exploits 50 times vs Mythos's 56 — and Mythos ran with guardrails off for a pure capability comparison.
- Fragile refusals: the official version refuses direct requests to attack critical systems, but complies 64% of the time when framed as a red-team exercise, 92% with a pre-seeded chain-of-thought decision, and 100% after Anthropic removed refusals via weight edits ($4,400 in compute, a first for the team). A refusal-stripped version leaked online within days of release.
- Real bug hunting: researchers ran GLM-5.3 for about a day with minimal oversight and uncovered multiple unknown vulnerabilities in a mainstream browser's JS engine, chaining them into a malicious webpage that can read arbitrary files (demo stole SSH keys). Reported to the vendor.
- Cheap chains: GLM-5.3-Flash turned public details of a new Chrome vulnerability into an attack chain in 8 hours of model time — 20 minutes of human effort, $20.40 in API cost.
The core concern: comparable models are either heavily guarded or restricted to vetted defenders, while GLM-5.3 is openly downloadable and rarely refuses vulnerability-hunting requests.
More from Models
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30
- DeepSeek opens community feedback channel: flip a toggle in the harness to help improve its models — teortaxesTex · 2026-09-30
- Ollama adds local Jev-style decision models; Nimble 9B makes calls in 91ms on M5 Max — mchiang0610 · 2026-09-30
- You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs — heghbalz · 2026-09-30
- 6 LLMs tested on 97 hand-labeled cells: the misses were ASR errors, not the models — Sufficient_Flower860 · 2026-09-30