DeepSeek V4.1 costs less to serve but prices 2x higher per output token, margins likely up

teortaxesTex · x · 2026-09-13

teortaxesTex argues V4.1 is strictly cheaper to serve than V4-Flash (lower FLOPs past 128K, 4x less cache, better caching) while priced 2x higher per output token even off-peak — up from 80% margins at V4's cheapest. A quoted post notes V4.1 Flash now scores above GPT-5.6 Luna Max on Artificial Analysis' intelligence index at low cost, making it a top API value pick, and pushing MiMo off the Pareto frontier.

Original post →

More from Models

Models channel →