DeepSeek V4.1 Finds 65.6% of CVEs on Cybersecurity Bench, Matches Frontier Models at a Fraction of the Cost
teortaxesTex · x · 2026-09-09
Third-party cybersecurity benchmarkers report DeepSeek V4.1 Flash is dramatically better: it rediscovers 65.6% of recent CVEs in a single run (up from 55.2%) and 84.4% at pass@3 (up from 75%), beating frontier models like Grok 4.6, Opus 5, and GPT-5.6-Sol. Precision also rose from 73.8% to 78.9% with fewer false positives, and per-task cost dropped thanks to more actions per turn and a 95.5% cache hit rate. Commentator teortaxesTex notes DeepSeek's consistent strength on this eval and argues that at this pace V4.1-Pro will surpass Astra at roughly 1/20th the cost — suggesting Astra may not be that large.
Related event: DeepSeek V4.1 Flash beta shows big gains in vision and cybersecurity(3 posts)→
More from Models
- VC communism is over: frontier models on rationing force hard model-choice thinking — StewartalsopIII · 2026-09-09
- Data labeling isn't dead: cherry-picked demos ≠ long-tail superiority — chris_j_paxton · 2026-09-09
- Apodex Open-Sources FrontierAgent: One-Command Local Agent Framework With Agent Team Mode — aakashgupta · 2026-09-09
- Apodex Claims 1.1 With Agent Team Sits in the Frontier Tier; 35B Mini Runs Locally — aakashgupta · 2026-09-09
- Apodex 1.1 Argues a Right Answer Can Still Be a Failed Task — aakashgupta · 2026-09-09
- User says they used 'gpt-6 astra' to build a visual essay on the Navier-Stokes problem — paw_lean · 2026-09-09