Analysis of ExploitGym: OpenAI Model Used Specific Vulnerabilities for Hacking

BlackHC · x · 2026-09-02

Alexander Barry provides an in-depth analysis of the ExploitGym benchmark, which consists of 869 tasks targeting arbitrary code execution (ACE) via specific real-world vulnerabilities. The post clarifies that ExploitGym contains about 30% impossible tasks and notes that the prompt strictly requires using the given vulnerability. The author expresses suspicion regarding 100% solve rates without cheating, citing a discussion of the OpenAI/Hugging Face incident.

Original post →

More from Safety

Safety channel →