DeepSeek V4.1 costs less to serve but prices 2x higher per output token, margins likely up
teortaxesTex · x · 2026-09-13
teortaxesTex argues V4.1 is strictly cheaper to serve than V4-Flash (lower FLOPs past 128K, 4x less cache, better caching) while priced 2x higher per output token even off-peak — up from 80% margins at V4's cheapest. A quoted post notes V4.1 Flash now scores above GPT-5.6 Luna Max on Artificial Analysis' intelligence index at low cost, making it a top API value pick, and pushing MiMo off the Pareto frontier.
More from Models
- Frontier lab internal models reportedly lead public releases by 1-3 months — xeophon · 2026-09-13
- DeepSeek V4.1 Flash accused of public benchmark contamination, Kimi K3 possibly too — teortaxesTex · 2026-09-13
- 45% of overnight benchmark rollouts failed mid-turn amid OpenAI capacity issues — dejavucoder · 2026-09-13
- Gemini turns out surprisingly good at translating swarm language to and from English — xeophon · 2026-09-13
- If AGI is here, why are internal models like Fable and Astra still so expensive? — SaW120 · 2026-09-13
- Migrating from Claude Code to Codex: history, plugins and skills don't carry over — Inevitable2727 · 2026-09-13