DeepSeek Bench Scores Rely on Max Reasoning, Inflating Task Time by Up to 10x

TheZachMueller · x · 2026-08-01

Community members note that DeepSeek's recently announced high benchmark scores were all achieved using max reasoning effort.

In this mode, the model consumes 2 to 10x more output tokens at a slightly lower drafter acceptance rate. Consequently, the actual wall time to complete tasks inflates to 5 to 10x compared to normal conditions. For instance, a task that typically takes 5 minutes could take 25 to 50 minutes under max reasoning, explaining why users might experience significant latency in practice.

Original post →

More from Models

Models channel →