DeepSeek Cybersecurity Test: Top Recall but Bottom Precision
teortaxesTex · x · 2026-08-13
In a recent cybersecurity vulnerability benchmark, DeepSeek demonstrated top-tier recall but suffered from severe precision issues.
- Top Recall: At pass@3, DeepSeek rediscovered 87.5% of benchmark CVEs, outperforming Opus 5 and Qwen 3.8.
- Bottom Precision: Only 65.6% of the vulnerabilities it reported were valid, far below GPT-5.6-Sol's 86.4%.
- Instability: The model's performance fluctuates significantly per run, averaging only 58.3% discovery rate, requiring multiple combined runs for optimal results.
A community developer argued that the unpredictable gains per run and high false positives indicate the model is still "under-post-trained."
More from Models
- Sakana Chat Major Update: New Models Namazu and Fugu Power Japanese Vibe Coding — SakanaAILabs · 2026-08-13
- Grok Offers 40% Off Extra Usage Credits, Up to $100 Discount — tetsuoai · 2026-08-13
- Yacine Declares Terminal Bench as the Only AI Benchmark That Matters Now — yacineMTB · 2026-08-13
- Opinion: Evaluation Harnesses Will Matter Less as Models Become More Capable — abacaj · 2026-08-13
- Grok-4.6 Gets Starstruck by a Low GitHub ID — nbaschez · 2026-08-13
- DeepSeek V4 Pro Finds RCE Vulnerability in Open-Source Project in Under 30 Minutes — teortaxesTex · 2026-08-13