GPT 6 Astra tops DeepSecBench, does defensive cyber tasks without refusals
cramforce · x · 2026-09-06
OpenAI's GPT 6 Astra takes the top spot on Vercel's DeepSecBench cybersecurity benchmark, displacing 5.6 Sol. Per-task cost is slightly higher, but benchmark runtime improves by over 2 hours. The key takeaway: the publicly available non-cyber version performs defensive security tasks without frequent refusals — a good time to re-run deepsec, an open-source full-repository security scanner that runs on your own infra with agent sandboxing.
Related event: GPT-6 Astra Tops Vercel's Security Benchmark at Half the Cost(2 posts)→
More from Models
- Models used to max benchmarks — now benchmarks are maxxing the model — abeirami · 2026-09-06
- GPT-6 Astra Builds a Full App From One Prompt in Under 10 Minutes — RileyRalmuto · 2026-09-06
- Zvi on Astra: model avoiding cheating because it'd get caught is actually worse — ZeroStateReflex · 2026-09-06
- Reddit post mocks community 'cope' as benchmarks get dismissed after Astra's flat reception — Tim_Apple_938 · 2026-09-06
- Scale puts AI capability and cost in a U-shape: GPT is absurdly cheap, Codex an unreal deal — chris_j_paxton · 2026-09-06
- Artificial Analysis overhauls Intelligence Index after GPT-6 Astra scoring skepticism — The Decoder · 2026-09-06