DeepSeek v4-flash: max thinking mode is both cheaper and faster than high
dosco · x · 2026-08-19
Developer dosco reports that with deepseek-v4-flash-latest, the "max" thinking mode is actually cheaper and faster than "high" mode for agentic workloads — a counterintuitive but practical data point for developers picking reasoning tiers.
More from Models
- GLM 5.3 scores 60 on AA overall intelligence index — zainhas · 2026-08-19
- GLM-5.3 tops AA Agentic Index, matching frontier models at a fraction of the cost — zainhas · 2026-08-19
- GLM 5.3 launches on ChatLLM with 30-day unlimited access — bindureddy · 2026-08-19
- DFlash2 on Qwen3.8 27B hits ~200tk/s for code, requires more VRAM — Hefty_Wolverine_553 · 2026-08-19
- Gemini Hallucinates Fictional Nodes When Assisting with ComfyUI — NickPassig · 2026-08-19
- Astra Delayed Again, Widening Gap Between OpenAI's Internal and Public Models — haider1 · 2026-08-19