DeepSeek V4-Flash Cuts Costs 100x, Sparking AI Economics Debate
The DeepSeek-V4-Flash-0731 model has recently sparked widespread discussion regarding the true inference costs of large models. According to Artificial Analysis, this model scored close to Fable 5 on the Terminal-Bench 2.1 evaluation, with a total cost for completing the same benchmark tasks being only 1/105 of the latter. This data has driven industry attention toward shifting AI economic metrics from "cost per token" to "cost per task completed," though it has also faced skepticism regarding its practical application stability and potentially misleading billing.
Confirmed
- In the Terminal-Bench 2.1 evaluation, DeepSeek V4-Flash scored 82.7, nearly matching Fable 5's score of 80.5.
- According to Artificial Analysis' test report, DeepSeek's overall cost to complete the same benchmark tasks is 105 times lower than Fable 5.
Unconfirmed
- API Stability and Actual Total Cost: In a 34-prompt test, developer @kmsdev found that the model's performance on OpenRouter providers was unstable, failing multiple generations and ultimately costing $1.29, losing to Kimi K3 in both performance and cost.
- Misleading Per-Token Pricing: @cephaloform and investor Chamath pointed out that although DeepSeek's per-token price is extremely low, if the model requires more rounds to complete a task, the actual total cost might be higher, meaning the low price could be misleading.
Why It Matters
- Shift in Evaluation Systems: Both @mustafamhus and @AravSrinivas noted that this two-order-of-magnitude cost reduction is exceptionally rare, marking DeepSeek's "2.0 moment." This requires enterprise decision-makers and developers to move beyond traditional per-token pricing mindsets and genuinely focus on "cost per task."
- Open-Source Model Comparison: In practical tests for specific tasks like game generation (@mustafamhus), despite DeepSeek V4-Flash's ultra-low cost, it still lagged behind Kimi K3 (9.5/10) and GLM 5.2 (9/10) in scoring, indicating that merely pursuing cost reduction cannot entirely replace the comprehensive capability competition among models.
2026-08-01 ~ 2026-08-03 · 11 related posts
Primary sources
- DeepSeek-V4-Flash Eval: Underperforms Kimi K3 in Both Cost and Score — kms_dev · 2026-08-01
- [source] DeepSeek V4-Flash completes benchmarks at 105x lower cost, but per-token price may mislead — cephaloform · 2026-08-02
- [source] DeepSeek V4-Flash Reportedly Achieves Same Benchmark at 105x Lower Cost — AravSrinivas · 2026-08-02
- DeepSeek Completes Same Benchmark Tasks at 105x Lower Cost Than Competitors — JosephJacks_ · 2026-08-02
- DeepSeek V4-Flash Completes Benchmarks at 1/105th the Cost of Fable — zainhas · 2026-08-02
- DeepSeek V4-Flash Rivals Fable 5: The AI War Shifts to Intelligence-Per-Dollar — mustafamhus · 2026-08-02
- [source] Kimi K3 vs GLM 5.2: DeepSeek V4 Flash is 150x Cheaper — mustafamhus · 2026-08-02
- DeepSeek V4 Flash Launch: AI Inference Costs 'Too Cheap to Meter' — intellectronica · 2026-08-02
- DeepSeek V4 Flash vs. OpenAI Luna: A Same-Tier Price and Performance Showdown — eyishazyer · 2026-08-02
- DeepSeek V4 Flash Matches Competitors in Coding at 1/34th the Cost — mattturck · 2026-08-03
1 near-duplicate retellings: mustafamhus