Re-estimating DeepSeek Inference Costs

teortaxesTex · x · 2026-07-20

Based on the DSpark paper and a previous V4 serving report, a post re-estimates the inference costs of the DeepSeek V4 series across different context lengths. The chart indicates: at 10K/100K/1M contexts, costs are roughly $0.764/$0.770/$0.821 per million output tokens for V4-Flash, and $2.870/$2.878/$2.952 for V4-Pro.

Using real-world production throughput and an assumption of $2/GPU-hour, the author estimates more realistic costs: V4-Flash at approx. $0.035/M output tokens, V4-Pro at $0.10/M, and a 50/50 mix at $0.067/M.

The author emphasizes that these production figures are much lower than batch-1 roofline estimates because continuous batching distributes MoE weight traffic across numerous concurrent requests. The 60–85% Flash and 57–78% Pro improvements mentioned primarily reflect latency/interactivity gains, not direct cost savings.

Original post →

More from Infra

Infra channel →