DeepSeek V4 Pro Subsequent Benchmark Retrospective

xeophon · x · 2026-07-19

This post summarizes retrospective findings on **DeepSeek V4 Pro**: the author collected 32 new benchmarks that emerged post-release, focusing on its performance against **GPT-5.4 / Claude Opus 4.6 / other frontier models**. The core conclusion is that the leading advantages touted during the initial release are unstable on subsequent benchmarks. In the comparisons shown, DeepSeek V4 Pro lags behind competitors in most areas, especially in agentic coding-related benchmarks. Based on this, the author questions whether the "best Chinese open-source model" narrative can survive until Kimi K3.

Original post →

More from Models

Models channel →