DeepSeek V4 Pro Subsequent Benchmark Retrospective
xeophon · x · 2026-07-19
This post summarizes retrospective findings on **DeepSeek V4 Pro**: the author collected 32 new benchmarks that emerged post-release, focusing on its performance against **GPT-5.4 / Claude Opus 4.6 / other frontier models**. The core conclusion is that the leading advantages touted during the initial release are unstable on subsequent benchmarks. In the comparisons shown, DeepSeek V4 Pro lags behind competitors in most areas, especially in agentic coding-related benchmarks. Based on this, the author questions whether the "best Chinese open-source model" narrative can survive until Kimi K3.
More from Models
- Grok 4.5 tops a long-horizon terminal benchmark, Elon Musk says — elonmusk · 2026-07-21
- K3 seems faster on the biggest coding plan than through OpenRouter, user says — xeophon · 2026-07-21
- Andrew Ng-style distillation joke turns model reuse into a theft punchline — pmddomingos · 2026-07-21
- Grok’s Twitter/X search quality appears to have regressed — ivan_bezdomny · 2026-07-21
- SuperGrok regresses on a simple X-account search task, user says — ivan_bezdomny · 2026-07-21
- Chinese LLM vendors push API prices lower as competition intensifies — sen_o · 2026-07-21