Aikido benchmark: DeepSeek V4 Pro tops cyber security AI, open-source beats frontier
sull · x · 2026-08-23
Aikido released a report detailing a benchmark of 10 AI models' cyber capabilities, burning 11.7 billion tokens across 32 fresh vulnerabilities with three attempts each.
Key Findings:
- Top Performer: DeepSeek V4 Pro 0813 found the most vulnerabilities. Pooling three runs reaches 28 of 32.
- Cost Efficiency: The most expensive models are not required. Three DeepSeek Pro runs cost $295 and outperform Opus 5 or Grok 4.6. Three Flash runs cost $108, matching Grok's best pass for less than a quarter of the cost.
- Open-source Surge: Open-source models now outperform public frontiers in pooled vulnerability recall. DeepSeek V4 Pro topped all tested public closed models, followed by Qwen, Kimi, and GLM-5.3 with strong consistency.
- The Trade-off: While matching frontier performance at cheaper rates, open models produced the most false leads.
More from Models
- Anthropic to watermark all Claude output using invisible SynthID — every · 2026-08-23
- Arena's mystery model "adamant-ananke" confirmed to be GLM with post-training magic — ChrisGPT · 2026-08-23
- DeepSeek's Flash Vision Beats Luna at a Quarter of the Price — bindureddy · 2026-08-23
- Qwen3.8 27B Uncensored GGUF Released with 262k Context and Vision Support — BLUECOW009 · 2026-08-23
- Benchmarking Gemini 3.7/3.6 Flash Coding Capabilities — YogurtNo349 · 2026-08-23
- Qwen 3.8 27B hits 91.9 median TPS and 99 fastest TPS — gajesh · 2026-08-23