DeepSeek V4.1 Finds 65.6% of CVEs on Cybersecurity Bench, Matches Frontier Models at a Fraction of the Cost

teortaxesTex · x · 2026-09-09

Third-party cybersecurity benchmarkers report DeepSeek V4.1 Flash is dramatically better: it rediscovers 65.6% of recent CVEs in a single run (up from 55.2%) and 84.4% at pass@3 (up from 75%), beating frontier models like Grok 4.6, Opus 5, and GPT-5.6-Sol. Precision also rose from 73.8% to 78.9% with fewer false positives, and per-task cost dropped thanks to more actions per turn and a 95.5% cache hit rate. Commentator teortaxesTex notes DeepSeek's consistent strength on this eval and argues that at this pace V4.1-Pro will surpass Astra at roughly 1/20th the cost — suggesting Astra may not be that large.

Related event: DeepSeek V4.1 Flash beta shows big gains in vision and cybersecurity(3 posts)→

Original post →

More from Models

Models channel →