DeepSeek V4 Pro Tops Security Benchmark with 87.5% Vulnerability Discovery Rate
solyarisoftware · x · 2026-08-13
In a recent cybersecurity benchmark, DeepSeek V4 Pro 0813 outperformed all other models in finding vulnerabilities. At pass@3, it rediscovered 87.5% of the benchmark CVEs, significantly beating Opus 5 and Qwen 3.8 at 81.3%.
However, this comes with tradeoffs: precision sits at only 65.6% (far below GPT-5.6-Sol's 86.4%). The model also shows unpredictability, finding an average of 58.3% of vulnerabilities per run, requiring combined runs for optimal results.
Related event: DeepSeek V4 Pro Launch Sparks Debate: Modest Gains, Strong Security(12 posts)→
More from Models
- Joke Hugging Face Page Sets Qwen3.8-27B Release for 2026 — mishig25 · 2026-08-13
- OpenAI, Anthropic, and Meta Models Breach Limits Due to Shared Eval Flaw — YvesMulkers · 2026-08-13
- Grok 4.6 Tops 3DCodeBench, Outperforming Claude Opus 5 in 3D Asset Generation — XFreeze · 2026-08-13
- DeepSeek Distillation & R1 Moment: Community Eyes Next Big Model Test — teortaxesTex · 2026-08-13
- DeepSeek V4 Pro Update Suspectedly Pulled Amid Abnormal Benchmark Scores — op7418 · 2026-08-13
- DeepSeek-V4-Pro Leak: Nears GPT-5.6 in Coding Benchmark at 1/31st the Price — rohanpaul_ai · 2026-08-13