Exploit Bench at 100%? New Model's Cybersecurity Score Stuns Developers

jarrodwatts · x · 2026-09-04

Developer Jarrod Watts flags a shock benchmark result: Exploit Bench apparently hit 100%, likely referring to GPT-6 Astra. The eval measures a model's ability to find and exploit vulnerabilities, and a perfect score would signal autonomous cyber capabilities approaching Preparedness Framework thresholds. Not yet officially confirmed.

Related event: GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6%(41 posts)→

Original post →

More from Models

Models channel →