DeepSeek V4 Pro Launch Sparks Debate: Modest Gains, Strong Security
DeepSeek released the V4 Pro 0813 version around August 13, but official silence on social media sparked user speculation. Multiple evaluations show that this version offers marginal performance improvements in standard benchmarks, almost identical to July's V4-Flash-0731, and even scores lower in some tests like SciCode, leading developers to complain about incremental updates or a flop. However, V4 Pro shines in cybersecurity and coding capabilities: it ranks second among open-source models (1607 points) in Chatbot Arena's WebDev coding arena, and fifth in open-source for the text arena (1465 points). In vulnerability mining benchmarks, it topped the list with an 87.5% recall rate, despite having the lowest precision. Additionally, its single-task cost is only $0.14, making it 17 times cheaper than the higher-ranked Kimi K3. User tests revealed that V4 Pro discovered an RCE vulnerability in an open-source project within 30 minutes, showing significantly stronger security capabilities than the Flash version; however, some developers found V4 Flash beating the Pro version in Canvas coding tasks. Overall, while V4 Pro excels in security and coding, its limited general performance improvements and the contrast between the low-key release and evaluation data have sparked community debate.
已确认
- 性能提升微弱:Multiple developers (like @teortaxesTex) noted that V4-0813 is nearly identical to V4-Flash-0731 in most benchmarks, with SciCode scores even lower.
- 代码能力开源第二:According to Chatbot Arena's AutoEval, DeepSeek-V4-Pro (Max) ranks roughly 8th overall and 2nd among open-source models in the WebDev coding arena with a score of 1607.
- 文本竞技场开源第五:Scored 1465 in AutoEval, ranking 5th in open-source, closely trailing GLM-5.1 (1467).
- 安全漏洞发现率登顶:Under the pass@3 setting, it rediscovered 87.5% of benchmark CVEs, higher than the 81.3% achieved by Opus 5 and Qwen 3.8.
- 成本优势:Single-task cost is $0.14, which is 17 times cheaper than Kimi K3.
- 官方沉默:DeepSeek did not heavily promote V4 Pro on social media, prompting user speculation.
尚未确认
- 性能提升是否真实:Some suggest V4 Pro might just be an optimized version of Flash, but official confirmation is lacking.
- 安全测试的精度问题:While recall is the highest, precision is at the bottom, and specific precision data was not provided in the materials.
- Vals 指数上涨 11 分:@teortaxesTex mentioned this score is inflated, but specific doubts were not detailed.
为什么重要
The release of DeepSeek V4 Pro has sparked discussions about incremental updates, but its breakthrough performance in security and coding, along with its extremely low cost, could alter the competitive landscape of open-source models. The contrast between the official low-profile attitude and community reviews also reflects the complexity of AI model evaluation and the gap in user expectations.
2026-08-13 ~ 2026-08-13 · 12 related posts
- Episode 1: DeepSeek V4 Pro Surfaces in Rumors, Imminent Release Expected(2026-08-11, 5 posts)
- Episode 2: DeepSeek V4 Specs and Pricing Allegedly Leaked(2026-08-11, 5 posts)
- Episode 3: DeepSeek V4 Pro 0813 Released with Major Agent Upgrades and Disruptive Pricing(2026-08-12, 27 posts)
- Episode 4: DeepSeek V4 Pro Launch Sparks Debate: Modest Gains, Strong Security(2026-08-13, 12 posts)
Primary sources
- DeepSeek v4-pro Beats Opus at Minimal Cost, But Real-World Differences Remain Marginal — haider1 · 2026-08-13
- DeepSeek V4-Pro Ranks #2 Open-Weight Model, Accused of Relying on pass@2 — teortaxesTex · 2026-08-13
- DeepSeek Cybersecurity Test: Top Recall but Bottom Precision — teortaxesTex · 2026-08-13
- DeepSeek V4 Pro Finds RCE Vulnerability in Open-Source Project in Under 30 Minutes — teortaxesTex · 2026-08-13
- [source] DeepSeek V4-0813 Shows No Gains Over July Version in Benchmarks — teortaxesTex · 2026-08-13
- DeepSeek V4-0813 Reported as Virtually Identical to Flash Version — teortaxesTex · 2026-08-13
- [source] DeepSeek-V4-Pro Ranks #2 Among Open Models in Code Arena — arena · 2026-08-13
- DeepSeek-V4-Pro Ranks #5 in Open Text Arena Category — arena · 2026-08-13
- DeepSeek V4 Pro Mocked for Scoring Only 1 Point Higher Than Flash — ns123abc · 2026-08-13
- [source] DeepSeek V4 Pro Tops Security Benchmark with 87.5% Vulnerability Discovery Rate — solyarisoftware · 2026-08-13
- Surprising Benchmark: DeepSeek V4 Flash Outperforms Pro in Coding — solyarisoftware · 2026-08-13
- DeepSeek V4 Pro Quietly Released: Intelligence Index 53, Above-Average Pricing — ns123abc · 2026-08-13