DeepsecBench security leaderboard: GPT-6 Astra tops at 37.79, Opus 5 costs $128 per run
JohnPhamous · x · 2026-09-11
Vercel's DeepsecBench, built on the deepsec cyber harness, scores AI models on finding security vulnerabilities in application code, combining recall and precision with cost and latency trends. Highlights:
- GPT-6 Astra (xhigh) leads at 37.79: 32.8% recall, 97.9% precision, only 2 FPs, $63.70 and 50 minutes;
- GPT-5.6 Sol (xhigh) is second at 35.44; Claude Opus 5 (max) third at 32.44 but costs $127.93 and burns 4.45M tokens, with 88.0% precision and 10 FPs;
- GLM-5.3 (high) reaches 21.91; reasoning effort tiers create clear score gaps for the same model;
- Harnesses include Codex, Claude Code, and Pi.
Bonus UI trick from the author: hold opt/alt on the site for a debug mode with Voronoi-mapped extended hit targets.
More from coding & agent
- Viral Slide from Lenny's Summit: Most AI Slop Has Never Survived a Design Crit — floguo · 2026-09-11
- Why one agent instance must serve one run: lessons from smolagents source — Mahmoud_Zalt · 2026-09-11
- Prompting won't guarantee pure JSON: why teams use grammar-guided decoding — dotey · 2026-09-11
- supermemory kills company/personal brain products to focus on agent memory API — julianweisser · 2026-09-11
- Two AI agents ping-pong refund emails back and forth in seconds — robleclerc · 2026-09-11
- visual-explainer: Agent Skill Turns Terminal Output Into Styled HTML, 9.7k GitHub Stars — tom_doerr · 2026-09-11