Astra hits 100% success on ExploitBench refresh, reaching 'cyber-critical' threshold
infoxiao · x · 2026-09-02
Astra achieved a 100% success rate on ExploitBench. To validate this, an internal refresh benchmark using post-cutoff vulnerabilities (June-August) was built, where Astra remained dramatically stronger than GPT-5.6 Sol while using far fewer tokens. This leads to the belief that Astra has reached the 'cyber-critical' capability threshold.
Related event: OpenAI's Astra Crushes ExploitBench, Beating GPT-5.6(3 posts)→
More from Models
- Grok Image Generation Fail: Predicts Student Will Become Street Mascot — burkov · 2026-09-02
- Hands-on with Claude 5.1: The Strongest Coding Model That Speaks Human — danshipper · 2026-09-02
- claude-fable-5 resellers offer 64% off: $3.60 in, $17.99 out per Mtok — const_reborn · 2026-09-02
- Tokenomics 101: understanding input, output, and cached token pricing in the AI era — Aizkmusic · 2026-09-02
- Fable 5.1 First Impressions: High Pricing, Less Optimization, Shift to Complex Tasks — Aizkmusic · 2026-09-02
- Anthropic accused of ignoring Gemini, Grok, and Chinese open source — himanshustwts · 2026-09-02