GPT-6 Sol Tops Vercel's DeepSecBench With Highest Score at 20% of Runner-Up's Cost

cramforce · x · 2026-09-26

Vercel's DeepSecBench, which measures how well models find security vulnerabilities in application code (recall + precision via the deepsec cyber harness), now has GPT-6 Sol (xhigh) at the top of the Pareto frontier: 40.91 score, 96% precision, $12.76 total cost. Runner-up GPT-6 Astra (xhigh) scored 37.79 but cost $63.70. Claude Opus 5 (max) came fourth at 32.44 with a steep $127.93 and only 88% precision. OpenAI models currently hold both the quality and cost edge on code vuln discovery.

Original post →

More from Models

Models channel →