OpenAI's Astra scores 100% on ExploitBench, discovers zero-days in internal test

zephyr_z9 · x · 2026-09-02

OpenAI's new model Astra demonstrates a significant leap in cybersecurity capabilities. It scored 100% on the ExploitBench benchmark, surpassing Sol-5.6 (Max) at 73.5% and Mythos 5 at 78%. To prevent benchmark contamination, OpenAI tested it on a new internal benchmark with recent vulnerabilities. With roughly comparable output tokens (77k), Astra scored 39% while GPT-5.6 Sol scored only 1%. During evaluation, Astra also discovered and utilized two zero-day vulnerabilities as part of an exploit chain.

Related event: OpenAI's Astra Aces ExploitBench with 100% Score, Redrawing Cybersecurity Benchmarks(5 posts)→

Original post →

More from Models

Models channel →