Anthropic shares Claude results on ExploitBench, with Opus 5 leading key safety metrics
TheZvi · x · 2026-07-25
Anthropic shares Claude results on ExploitBench
The post highlights Anthropic’s latest results on ExploitBench, a benchmark for measuring how models handle exploit-style tasks and sandbox-escape attempts.
The figure shared in the post compares several Claude models on the benchmark and shows:
- Claude Opus 5 with the highest AutoNudge Mean among the listed Claude models
- Claude Opus 5 also leading in AutoNudge Cap% and Full ACEs
- Older or smaller Claude variants scoring lower across the same metrics
The benchmark description says the evaluation was run across multiple trials and environments, and that Full ACEs refers to complete exploits that achieve arbitrary code execution combined across both plain and AutoNudge settings.
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11