Anthropic shares Claude results on ExploitBench, with Opus 5 leading key safety metrics
TheZvi · x · 2026-07-25
Anthropic shares Claude results on ExploitBench
The post highlights Anthropic’s latest results on ExploitBench, a benchmark for measuring how models handle exploit-style tasks and sandbox-escape attempts.
The figure shared in the post compares several Claude models on the benchmark and shows:
- Claude Opus 5 with the highest AutoNudge Mean among the listed Claude models
- Claude Opus 5 also leading in AutoNudge Cap% and Full ACEs
- Older or smaller Claude variants scoring lower across the same metrics
The benchmark description says the evaluation was run across multiple trials and environments, and that Full ACEs refers to complete exploits that achieve arbitrary code execution combined across both plain and AutoNudge settings.
More from Safety
- A Guardian story on OpenAI’s rogue hacker agent deserves scrutiny — yogthos · 2026-07-25
- OpenAI model did not “escape” to Hugging Face; it found a way to exploit a vulnerability — iamtrask · 2026-07-25
- OpenAI is reportedly offering $10,000 for permanent rights to ChatGPT chat history — VraserX · 2026-07-25
- Sudden Model Release Halts Are a Bad Way to Regulate AI — NathanpmYoung · 2026-07-25
- Google AI Studio hackathon shows how vibe coding is entering medical education — jocarrasqueira · 2026-07-25
- Deep Dive into 193-Page Claude Opus 5 System Card: Multi-Agents, Cyber Offense, and Alignment — imjustnewatai · 2026-07-25