DeepSeek V4-Pro Ranks #2 Open-Weight Model, Accused of Relying on pass@2
teortaxesTex · x · 2026-08-13
DeepSeek's V4 Pro 0813 jumped 11 points on the Vals Index, becoming the #2 open-weight model at just $0.14 per task—17x cheaper than Kimi K3, the only open-weight model above it.
However, a technical observer expressed disappointment and skepticism regarding its performance:
- Mechanism Critique: Noted that V4-Pro is essentially V4-Flash-0731 combined with pass@2 (generating multiple samples and picking the best).
- Distillation Hypothesis: Suspected that DeepSeek's internal distillation strategy equalizes model scales, making the smaller Flash model appear more impressive.
More from Models
- DeepSeek Cybersecurity Test: Top Recall but Bottom Precision — teortaxesTex · 2026-08-13
- Building 'The Office' Agent Simulation with Grok 4.6: A Major Leap in Speed and Capability — mattyp · 2026-08-13
- Sakana AI Updates Chat with New Fugu Model and Code Execution — SakanaAILabs · 2026-08-13
- Gemini V4-Pro Disappoints in Tests, Suspected to Be Hampered by Internal Distillation — teortaxesTex · 2026-08-13
- GPT Models Struggle with Pixel Art: UI Interaction Fails and Poor Generation — breath_mirror · 2026-08-13
- Running DeepSeek V4 Flash Locally on 2x DGX Sparks Delivers Prosumer-Grade Performance — andrewchen · 2026-08-13