Re-estimating DeepSeek Inference Costs
teortaxesTex · x · 2026-07-20
Based on the DSpark paper and a previous V4 serving report, a post re-estimates the inference costs of the DeepSeek V4 series across different context lengths. The chart indicates: at 10K/100K/1M contexts, costs are roughly $0.764/$0.770/$0.821 per million output tokens for V4-Flash, and $2.870/$2.878/$2.952 for V4-Pro.
Using real-world production throughput and an assumption of $2/GPU-hour, the author estimates more realistic costs: V4-Flash at approx. $0.035/M output tokens, V4-Pro at $0.10/M, and a 50/50 mix at $0.067/M.
The author emphasizes that these production figures are much lower than batch-1 roofline estimates because continuous batching distributes MoE weight traffic across numerous concurrent requests. The 60–85% Flash and 57–78% Pro improvements mentioned primarily reflect latency/interactivity gains, not direct cost savings.
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11