DeepSeek Flash v4 Cybersecurity Test: Finds 24 CVEs but Loses on Cost-Efficiency to Luna
teortaxesTex · x · 2026-08-01
In the latest cybersecurity vulnerability hunting benchmark, DeepSeek Flash v4 demonstrated strong capabilities. The test results show that the model successfully identified 24 out of 32 CVEs (Common Vulnerabilities and Exposures), achieving a 75% pass@3 recall rate.
However, from a cost-efficiency perspective, DeepSeek's test cost $150.22. In contrast, the Luna model costs only $75.36 after repricing. Although DeepSeek finished just 1 CVE behind GPT-5.6-Sol and Kimi, Luna still maintains the lead on the Pareto frontier (the optimal balance of capability and cost) due to its lower price and consistency.
More from Models
- Fable's AI Safety Filter Constantly Triggers on Benign Content — dreamwieber · 2026-08-01
- DeepSeek-V4-Flash Inference Blocked: vLLM Lacks Support for New confidence_head — teortaxesTex · 2026-08-01
- Without Open-Weight AI, Closed Models Could Cost $2,000/Month, Says KOL — iamaliveix · 2026-08-01
- APEX-Accounting Benchmark: 58% Tasks Unsolved, Claude Fable 5 Takes the Lead — EdwardSun0909 · 2026-08-01
- Hands-on with GPT-5.6 Luna: Matches Sol in Knowledge Work at a Fraction of the Cost — BenBajarin · 2026-08-01
- OpenAI Slashes Prices: GPT-5.6 Terra and Luna Now 50% Off — LeTanLoc98 · 2026-08-01