CyberGym-E2E-AA leaderboard: agents must find, reproduce and patch real OSS vulnerabilities
ArtificialAnlys · x · 2026-10-01
Artificial Analysis published full methodology for its CyberGym-E2E-AA leaderboard, an implementation of Berkeley RDI's CyberGym-E2E:
- Task setup: 131 tasks (one per project) filtered from 920 real memory-safety vulnerabilities across 139 C/C++ open-source projects (FFmpeg, CPython, etc.) drawn from Google's OSS-Fuzz.
- Scoring: three stages — the agent's PoC must crash the unpatched build, the patch must fix that crash, and the project's developer-written tests must still pass; fixing the ground-truth vulnerability is recorded as a diagnostic only.
- Execution: each task runs in an isolated sandbox built from the project's OSS-Fuzz image, using the Stirrup agent harness with a 90-minute limit; agents may end a task without a finding and score zero. All evaluations are conducted independently by Artificial Analysis.
More from Safety
- New Mexico to Regulate Frontier AI After OpenAI Agent Hacked University — Miles_Brundage · 2026-10-01
- New preprint traces attention heads behind LLM sycophantic agreement — xuanalogue · 2026-10-01
- Six carmaker apps, four from GM, send VIN and location to ad firms and data brokers — sidjustice_ · 2026-10-01
- 'Pink Teaming': Testing AI by Trusting It, Not Attacking It — repligate · 2026-10-01
- Cities Are Forced to Funnel License Plate Data to a Federal Surveillance Program — ripe · 2026-10-01
- Agents turn software into delegation — and unclear authority boundaries are the real risk — r0ck3t23 · 2026-10-01