DeepSeek Bench Scores Rely on Max Reasoning, Inflating Task Time by Up to 10x
TheZachMueller · x · 2026-08-01
Community members note that DeepSeek's recently announced high benchmark scores were all achieved using max reasoning effort.
In this mode, the model consumes 2 to 10x more output tokens at a slightly lower drafter acceptance rate. Consequently, the actual wall time to complete tasks inflates to 5 to 10x compared to normal conditions. For instance, a task that typically takes 5 minutes could take 25 to 50 minutes under max reasoning, explaining why users might experience significant latency in practice.
More from Models
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Sakana AI translation outperforms Google and DeepL in Japanese-English benchmarks — SakanaAILabs · 2026-08-24
- Developer haider makes his own LLM tier list after disagreeing with theo's rankings — haider1 · 2026-08-24
- Mystery OxAlpha Beats Claude; Alibaba Raises $10B for AI — 创业邦 · 2026-08-24
- OpenAI and Google cut LLM prices; mystery OxAlpha model beats Claude on DeepSWE — 创业邦 · 2026-08-24
- AI News Digest: DeepSeek Weekend Discounts, GPT-5.6 Sol Price Cut, Alibaba's $10B AI Raise — APPSO · 2026-08-24