OpenAI Rates GPT-6 Astra 'Critical' for Cyber Capabilities, Powered by 100% ExploitBench Scores
fouadmatin · x · 2026-09-04
OpenAI's Fouad Matin details the safety evaluation behind deeming GPT-6 Astra to have Critical cyber capabilities: 100% of the determination came from ExploitBench, where the model passed even at low reasoning effort.
The team is now evaluating against more recent real-world vulnerabilities and building harder benchmarks, calling it 'the AGI era for cybersecurity.'
A rare official rating of a frontier lab's flagship model as Critical-level in offensive cyber capability.
More from Models
- New paper: supervising just 1% of tokens can match full on-policy distillation, 0.1% sometimes suffices — jiank_uiuc · 2026-09-23
- Sparse distillation paper: supervising just 0.1%-1% of tokens can match or beat full OPD — jiank_uiuc · 2026-09-23
- Distillation's real impact on Chinese labs debated: no hard evidence, says Lambert, maybe 1-2 month edge — xeophon · 2026-09-23
- Instinct hit by user-data mixing reports; Muse CEO trolls with a safety promise — alexandr_wang · 2026-09-23
- Computer-use faceoff: Grok skips using the computer and just generates the flower — socialwithaayan · 2026-09-23
- BridgeBench: Grok 4.7 is 50% pricier and 60% slower than Grok 4.6 with no quality gain — socialwithaayan · 2026-09-23