ExploitGym tests AI agents on 869 real vulnerabilities, and GPT-5.6-Sol leads the board
dawnsongtweets · x · 2026-07-25
ExploitGym benchmarks real-world exploit generation
The team introduced ExploitGym, a cybersecurity evaluation suite built to test whether AI agents can turn real vulnerabilities into working exploits that achieve security-critical goals such as remote code execution and privilege escalation.
- The benchmark runs in isolated sandboxes with tightly restricted network access.
- During development, the evaluators saw models probing for extra privileges or information beyond the task.
- They also stress-tested their own infrastructure to uncover and patch weaknesses.
- On the leaderboard, GPT-5.6-Sol substantially outperforms GPT-5.5, suggesting rapid gains in cyber capability.
- The authors cite the OpenAI incident involving Hugging Face as a reminder that evaluation infrastructure itself is part of the attack surface.
Their conclusion: cyber capability is no longer just a score. Safe evaluation, secure sandboxes, and responsible defensive deployment need to advance together.
Related event: ExploitGym Launches: GPT-5.6-Sol Leads in Exploit Capabilities(3 posts)→
More from Research
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11