Testing DeepSeek-V4-Flash: 'Low' Reasoning Effort Generates More Tokens Than 'High'
coder543 · reddit · 2026-08-03
A Reddit user tested the four reasoning effort modes of DeepSeek-V4-Flash-0731 (None, Low, High, Max) and discovered a counterintuitive behavior: the Low mode generates more tokens than the High mode.
The author validated this across both a local quantized version and the official DeepSeek API (averaging 20 requests per mode):
- Low mode: Averaged around 1,200-1,300 total tokens, with 800-900 tokens spent on the reasoning process alone.
- High mode: Averaged only 500-600 total tokens, with much more concise reasoning and final answers.
- Max mode: Highly verbose locally, but more constrained on the API.
The author suggests that DeepSeek and benchmarking platforms like Artificial Analysis should publish results for all effort modes, rather than just Max. They also noted a current bug on OpenRouter that breaks the reasoning effort modes.
More from Models
- Researcher: GPT 5.6 Sol Ultra Beats Pro for Long-Horizon Hard Problems — arankomatsuzaki · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24