DeepSeek v4 Pro Review: Matches Flash on Most Evals, Raises Scaling Questions

teortaxesTex · x · 2026-08-17

Community evaluations reveal that DeepSeek v4 Pro (0813) scores 66.2% on the WeirdML benchmark. While it shows solid improvement over the initial v4 Pro and slightly leads v4 Flash (63.0%), it performs similarly to Flash on most other evaluations. Observers suggest this might indicate limited scalability of the v4 architecture.

Original post →

More from Models

Models channel →