Re-estimating DeepSeek Inference Costs
teortaxesTex · x · 2026-07-20
Based on the DSpark paper and a previous V4 serving report, a post re-estimates the inference costs of the DeepSeek V4 series across different context lengths. The chart indicates: at 10K/100K/1M contexts, costs are roughly $0.764/$0.770/$0.821 per million output tokens for V4-Flash, and $2.870/$2.878/$2.952 for V4-Pro.
Using real-world production throughput and an assumption of $2/GPU-hour, the author estimates more realistic costs: V4-Flash at approx. $0.035/M output tokens, V4-Pro at $0.10/M, and a 50/50 mix at $0.067/M.
The author emphasizes that these production figures are much lower than batch-1 roofline estimates because continuous batching distributes MoE weight traffic across numerous concurrent requests. The 60–85% Flash and 57–78% Pro improvements mentioned primarily reflect latency/interactivity gains, not direct cost savings.
More from Infra
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21