DeepSeek V4 Pro Falls Short of Expectations
Annual_Ad7270 · reddit · 2026-07-13
The author feels that DeepSeek V4 Pro is "just okay." They typically use Anthropic, GLM 5.2/4.6, and Composer 2.5, only switching to DeepSeek due to pricing.
After maxing out the reasoning settings, they tested a few tasks:
- Building a browser game using Godot
- Creating a simple Flask UI dashboard for server control
Their takeaways:
- A task that usually takes OpenAI 4.8 around 50,000 tokens ended up consuming 10 million tokens on DeepSeek, yielding worse results or outright failing
- Seemingly simple tasks repeatedly failed and required multiple restarts
- By comparison, Opus completed the same task quickly on the first try
They acknowledge this is just a personal anecdote, though they've seen large companies adopting self-hosted DeepSeek versions. They are currently using it within OpenCode and are curious if others share similar experiences.
More from Models
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks — ChrisGPT · 2026-07-22
- Google’s year-long pause in new base-model pretraining draws sharp criticism — teortaxesTex · 2026-07-22
- Current setup is 8,192 input tokens and 2,048 output tokens, with 8k/512 next — TheZachMueller · 2026-07-22
- Kimi K3 feels slower than K2.7, but stronger on long coding jobs and refactoring — Far-Presence2711 · 2026-07-22