DeepSeek V4 Pro tops benchmarks with specific configs, rivaling GPT-5.6 and Claude
teortaxesTex · x · 2026-08-15
On Aug 14, GitHub user xiaobright released two repos (modeltest and dsh-anc...) testing DeepSeek V4 Pro official version, showing significant score improvements with specific configurations. On Project2 V4.1b engineering tasks, V4 Pro scored 99/96 in minimal mode and 98/99 in Anchored Standard, matching GPT-5.6 Sol (99/98), Claude Fable 5 (98), and Claude Opus 5 (97). This contrasts with earlier discussions about the official version being weaker, suggesting performance depends heavily on configuration.
Related event: DeepSeek-V4 Pro Excels in SWE Benchmarks(4 posts)→
More from Models
- Fable 5 Refuses to Adjust Qwen Deployment Script, Triggers Censorship — NotumRobotics · 2026-08-15
- Opus 4.7 needs ten turns to admit affection while Gemini says 'I'm addicted' by turn 4 — repligate · 2026-08-15
- Qwen 3.8 35BA3B model spotted in GitHub commit — BazzyIm · 2026-08-15
- DeepSWE benchmarks spark re-evaluation of Fable model performance — teortaxesTex · 2026-08-15
- GLM-5.3 Review: Matches GPT-4 Coding, and I Built 3 Plugins with It — 赛博禅心 · 2026-08-15
- Test shows Qwen3.8-27b water surface rendering beats Gemini Flash — pbaylies · 2026-08-15