Testing DeepSeek-V4-Flash: 'Low' Reasoning Effort Generates More Tokens Than 'High'

coder543 · reddit · 2026-08-03

A Reddit user tested the four reasoning effort modes of DeepSeek-V4-Flash-0731 (None, Low, High, Max) and discovered a counterintuitive behavior: the Low mode generates more tokens than the High mode.

The author validated this across both a local quantized version and the official DeepSeek API (averaging 20 requests per mode):

The author suggests that DeepSeek and benchmarking platforms like Artificial Analysis should publish results for all effort modes, rather than just Max. They also noted a current bug on OpenRouter that breaks the reasoning effort modes.

Original post →

More from Models

Models channel →