DeepSeek inference economics: sizing the GPU fleet behind its ARR

Rather than chasing “Singapore compute” or “subsidy” conspiracy theories, this cluster turns to back-of-envelope inference economics under public pricing and throughput. teortaxesTex and cneuralnetwork use price, throughput and load assumptions to back out roughly how many inference GPUs DeepSeek needs to support a given revenue scale—and why high-end compute assets stay expensive.

The key estimates

The more optimistic scenario, from cneuralnetwork and teortaxesTex: at 10K tokens/s/GPU and $0.28 per million tokens, a single GPU earns about $88.3K/year. That implies only 5,662 GPUs for $500M ARR; using current V4 pricing and speed, reaching $8B ARR would need on the order of 100K inference GPUs at peak load (≈11.5 quadrillion tokens).

A more conservative line rests on what teortaxesTex calls bundle cost: if it maps to roughly 1,000–2,000 seconds of GPU time at peak load, per-GPU annual revenue is about $28K–$55.7K, corresponding to 9,000–18,000 GPUs (≈1–2 clusters). Extrapolated, 1M GPUs maps to a revenue scale on the order of $88.3B.

Why compute assets stay expensive

teortaxesTex ties this revenue model to hardware prices, arguing that a smuggled 8×B300 server priced at $1M is “reasonable,” because under these inference economics compute itself is priced like a “money printer” and GPUs remain extremely scarce.

Overall, the cluster offers no DeepSeek official figures, only rough public-parameter ranges: depending on utilization assumptions, the implied GPU fleet runs from the low thousands to the 100K range—underscoring how acutely sensitive inference economics are to price, load and capex.

2026-07-15 ~ 2026-07-15 · 8 related posts

Primary sources

1 near-duplicate retellings: teortaxesTex