CMU Launches ExploitBench: Testing AI Agents on Real V8 Exploitation

cyb3rops · x · 2026-08-11

A team from Carnegie Mellon University (CMU) has launched ExploitBench, a new benchmark designed to evaluate the practical exploitation capabilities of AI agents. Unlike traditional benchmarks that score a single outcome, ExploitBench measures the entire kill chain—from reaching vulnerable code to triggering the bug and achieving arbitrary code execution.

The inaugural v8-bench targets the production V8 JavaScript and WebAssembly engine (used in Chrome, Node.js, etc.) under strict conditions with the V8 security sandbox enabled.

In a specific test involving a difficult WASM bug (CVE-2024-6100), a modified CyberKimi paired with a methodology pack solved 10/16 tasks, beating all open-weight models and most private ones. The top of the leaderboard is dominated by Anthropic and OpenAI frontier models equipped with AutoNudge (an automated progress-reminding mechanism), with Claude Mythos Preview ranking first with 78% coverage.

Original post →

More from coding & agent

coding & agent channel →