DeepSeek v4 Pro Eval Sparks Debate: Cherry-picked Benchmarks Obscure True Power
kristoph · x · 2026-08-13
A user compared DeepSeek v4 Pro and Grok 4.6, noting that DeepSeek costs about 1/7th of Grok. However, a reply criticized AI labs for cherry-picking benchmarks, making it impossible to tell how good DeepSeek v4 Pro actually is compared to Grok 4.6.
Related event: DeepSeek V4 Pro vs. Grok 4.6: Performance and Cost Controversy(2 posts)→
More from Models
- Grok 4.6 tested on bug bench: outperforms predecessor, becomes new default — PawelHuryn · 2026-08-13
- Grok 4.6 tops AI models under $5 workload budget, showing high intelligence per dollar — XFreeze · 2026-08-13
- OpenAI and Anthropic Models Dominate in Long-Running Autonomous Workflows — scaling01 · 2026-08-13
- Grok 4.6 Nearly Matches Claude Fable 5 on Agentic Benchmark at a Fraction of the Cost — ArtificialAnlys · 2026-08-13
- DeepSeek V4 Pro Offers 10x Cheaper Cost Per Task Than GLM5.2 — ojasvi_yadav · 2026-08-13
- DeepSeek API Requests Timeout Amid Suspected Technical Issues — cedric_chee · 2026-08-13