DeepSeek V4 Pro tops benchmarks with specific configs, rivaling GPT-5.6 and Claude

teortaxesTex · x · 2026-08-15

On Aug 14, GitHub user xiaobright released two repos (modeltest and dsh-anc...) testing DeepSeek V4 Pro official version, showing significant score improvements with specific configurations. On Project2 V4.1b engineering tasks, V4 Pro scored 99/96 in minimal mode and 98/99 in Anchored Standard, matching GPT-5.6 Sol (99/98), Claude Fable 5 (98), and Claude Opus 5 (97). This contrasts with earlier discussions about the official version being weaker, suggesting performance depends heavily on configuration.

Related event: DeepSeek-V4 Pro Excels in SWE Benchmarks(4 posts)→

Original post →

More from Models

Models channel →