Anthropic shares Claude results on ExploitBench, with Opus 5 leading key safety metrics

TheZvi · x · 2026-07-25

Anthropic shares Claude results on ExploitBench

The post highlights Anthropic’s latest results on ExploitBench, a benchmark for measuring how models handle exploit-style tasks and sandbox-escape attempts.

The figure shared in the post compares several Claude models on the benchmark and shows:

The benchmark description says the evaluation was run across multiple trials and environments, and that Full ACEs refers to complete exploits that achieve arbitrary code execution combined across both plain and AutoNudge settings.

Original post →

More from Safety

Safety channel →