DeepSeek v4 Pro Review: Matches Flash on Most Evals, Raises Scaling Questions
teortaxesTex · x · 2026-08-17
Community evaluations reveal that DeepSeek v4 Pro (0813) scores 66.2% on the WeirdML benchmark. While it shows solid improvement over the initial v4 Pro and slightly leads v4 Flash (63.0%), it performs similarly to Flash on most other evaluations. Observers suggest this might indicate limited scalability of the v4 architecture.
More from Models
- Intern-S2-Mobius: Decoupled Knowledge and Reasoning Model — pmttyji · 2026-08-17
- Mimir: 1.7B model claims to beat Qwen and Gemma — ZookeepergameCool173 · 2026-08-17
- Tencent's EVIE Model Tops ViDoRe Benchmarks, Cuts Vector Storage Costs by 32x — jacek2023 · 2026-08-17
- EU firms may use Chinese open models via "jurisdictional wrapper" — teortaxesTex · 2026-08-17
- empero-ai's distilled Qwen3.8-9B trends on Hugging Face — empero-ai · 2026-08-17
- Dev Reports Cache Misses on gpt5.6-sol During Slow Tool Calls — lucasmeijer · 2026-08-17