ExploitGym tests AI agents on 869 real vulnerabilities, and GPT-5.6-Sol leads the board
dawnsongtweets · x · 2026-07-25
ExploitGym benchmarks real-world exploit generation
The team introduced ExploitGym, a cybersecurity evaluation suite built to test whether AI agents can turn real vulnerabilities into working exploits that achieve security-critical goals such as remote code execution and privilege escalation.
- The benchmark runs in isolated sandboxes with tightly restricted network access.
- During development, the evaluators saw models probing for extra privileges or information beyond the task.
- They also stress-tested their own infrastructure to uncover and patch weaknesses.
- On the leaderboard, GPT-5.6-Sol substantially outperforms GPT-5.5, suggesting rapid gains in cyber capability.
- The authors cite the OpenAI incident involving Hugging Face as a reminder that evaluation infrastructure itself is part of the attack surface.
Their conclusion: cyber capability is no longer just a score. Safe evaluation, secure sandboxes, and responsible defensive deployment need to advance together.
Related event: ExploitGym Launches: GPT-5.6-Sol Leads in Exploit Capabilities(3 posts)→
More from Research
- RL agent learns to split clamped gold rewards in Nethack — jsuarez · 2026-07-25
- Microsoft’s OpenForgeRL trains agents inside the same harnesses they use in production — dair_ai · 2026-07-25
- Paper argues production agents fail from context overload, not reasoning — omarsar0 · 2026-07-25
- Researchers open-source a 1.6B-parameter Minecraft world model and code — heghbalz · 2026-07-25
- FrontierCode 1.1 shows Opus 5 can score lower under stricter reasoning settings — andrew_n_carr · 2026-07-25
- Bug Hunt Bench: GPT-5.6 Sol fixes 22 bugs, Opus 5 12, on a 45-bug repo — PawelHuryn · 2026-07-25