GPT-6 Sol Tops Vercel's DeepSecBench With Highest Score at 20% of Runner-Up's Cost
cramforce · x · 2026-09-26
Vercel's DeepSecBench, which measures how well models find security vulnerabilities in application code (recall + precision via the deepsec cyber harness), now has GPT-6 Sol (xhigh) at the top of the Pareto frontier: 40.91 score, 96% precision, $12.76 total cost. Runner-up GPT-6 Astra (xhigh) scored 37.79 but cost $63.70. Claude Opus 5 (max) came fourth at 32.44 with a steep $127.93 and only 88% precision. OpenAI models currently hold both the quality and cost edge on code vuln discovery.
More from Models
- "Local Models Are Useless" Sparks Pushback From a User Who Cut All AI Subs — Merzmensch · 2026-09-26
- Opus 5.5 almost entirely stopped using em dashes, users observe — Outside-Iron-8242 · 2026-09-26
- Interfaze AI Open-Sources Lev, a 4B System-One Model Built on Qwen — ycombinator · 2026-09-26
- Hugo Bowne-Anderson: claims that Opus 5.5 writes better prose are nonsense — hugobowne · 2026-09-26
- Alibaba Claims RSI Progress: Qwen3.8-Max Self-Improved Its AA Score From 40 to 45 Over 33 Autonomous Cycles — teortaxesTex · 2026-09-26
- Opus 5.5 unblocks a 4-month ts-rust port in 10 hours where GPT-6 Astra stalled at 85% — EricBuess · 2026-09-26