CMU Launches ExploitBench: Testing AI Agents on Real V8 Exploitation
cyb3rops · x · 2026-08-11
A team from Carnegie Mellon University (CMU) has launched ExploitBench, a new benchmark designed to evaluate the practical exploitation capabilities of AI agents. Unlike traditional benchmarks that score a single outcome, ExploitBench measures the entire kill chain—from reaching vulnerable code to triggering the bug and achieving arbitrary code execution.
The inaugural v8-bench targets the production V8 JavaScript and WebAssembly engine (used in Chrome, Node.js, etc.) under strict conditions with the V8 security sandbox enabled.
In a specific test involving a difficult WASM bug (CVE-2024-6100), a modified CyberKimi paired with a methodology pack solved 10/16 tasks, beating all open-weight models and most private ones. The top of the leaderboard is dominated by Anthropic and OpenAI frontier models equipped with AutoNudge (an automated progress-reminding mechanism), with Claude Mythos Preview ranking first with 78% coverage.
More from coding & agent
- Pydantic AI Harness v0.18.1 Released — solyarisoftware · 2026-08-11
- LlamaIndex's Jerry Liu Launches LiteParse: 200-Page Docs in 4ms — solyarisoftware · 2026-08-11
- Building AI Agents: Why External Systems Offer the Best Memory — goyalshaliniuk · 2026-08-11
- A Glance at the 7 Types of Memory in AI Agents — goyalshaliniuk · 2026-08-11
- 7 Types of Memory Systems for Production AI Agents — goyalshaliniuk · 2026-08-11
- Multi-Agent Tool for Automated Company Research Hits 2.2k Stars — tom_doerr · 2026-08-11