DeepsecBench vulnerability-finding eval: GPT-6 Sol leads at 46%
ArtificialAnlys · x · 2026-09-28
Artificial Analysis highlights DeepsecBench-AA (from Vercel), which isolates vulnerability discovery: given a codebase and a budget, an agent must find every vulnerability, scored by F2 against an expert-verified golden set, weighting recall over precision.
- GPT-6 Sol (max) leads at 46%, ahead of GPT-6 Astra (37%), DeepSeek V4.1 Flash (31%), Grok 4.7 (27%) and Claude Opus 5.5 (27%)
- The Score-vs-Cost Pareto frontier: GPT-6 Sol at $1.83/task, DeepSeek V4.1 Flash at $0.31/task, MiMo-V2.6-Pro at $0.09/task
More from Models
- Early Sonnet 5.5 impressions: ~5x faster bug fixes at high effort — burkov · 2026-09-29
- Why 4o feels different: thread argues native omni training, not capability, shapes model personality — RileyRalmuto · 2026-09-29
- Sonnet 5.5 effort settings make no difference in 15-task coding test: 9/15 at low, medium and high — every · 2026-09-29
- PrunaAI claims its text-to-video modes sit on DesignArena Pareto frontiers — guennemann · 2026-09-29
- Reddit user: Opus 5.5 silently falls back to Opus 5 on nearly every prompt — fishcat_catfish · 2026-09-29
- Sonnet 5.5 clones open-source editor Proof at low effort, joining elite group of just four models — every · 2026-09-29