DeepSeek inference economics: sizing the GPU fleet behind its ARR

Rather than chasing “Singapore compute” or “subsidy” conspiracy theories, this cluster turns to back-of-envelope inference economics under public pricing and throughput. teortaxesTex and cneuralnetwork use price, throughput and load assumptions to back out roughly how many inference GPUs DeepSeek needs to support a given revenue scale—and why high-end compute assets stay expensive.

The key estimates

The more optimistic scenario, from cneuralnetwork and teortaxesTex: at **10K tokens/s/GPU** and **$0.28 per million tokens**, a single GPU earns about **$88.3K/year**. That implies only ~**5,662 GPUs** for **$500M ARR**; using current V4 pricing and speed, reaching **$8B ARR** would need on the order of **100K** inference GPUs at peak load (≈11.5 quadrillion tokens).

A more conservative line rests on what teortaxesTex calls bundle cost: if it maps to roughly **1,000–2,000 seconds** of GPU time at peak load, per-GPU annual revenue is about **$28K–$55.7K**, corresponding to **9,000–18,000** GPUs (≈1–2 clusters). Extrapolated, **1M GPUs** maps to a revenue scale on the order of **$88.3B**.

Why compute assets stay expensive

teortaxesTex ties this revenue model to hardware prices, arguing that a smuggled **8×B300** server priced at **$1M** is “reasonable,” because under these inference economics compute itself is priced like a “money printer” and GPUs remain extremely scarce.

Overall, the cluster offers no DeepSeek official figures, only rough public-parameter ranges: depending on utilization assumptions, the implied GPU fleet runs from the low thousands to the 100K range—underscoring how acutely sensitive inference economics are to price, load and capex.

2026-07-15 ~ 2026-07-15 · 8 related posts

1 near-duplicate retellings: teortaxesTex