DeepSeek post pairs inference-system optimizations with a 545% theoretical margin analysis
teortaxesTex · x · 2026-07-22
A repost about DeepSeek combines an inference-system overview with a separate analysis of its theoretical revenue and expenses. The screenshot cites throughput optimizations such as cross-node EP batch scaling, computation-communication overlap, and load balancing, plus claims about H800 token throughput and a theoretical 545% cost profit margin.
The attached chart and paper page also compare network and latency choices such as RoCE vs. InfiniBand, discussing congestion control, routing, and GPU-direct communication tradeoffs for large-scale AI workloads.
More from Infra
- BloombergNEF: U.S. data centers could use 20% of national electricity by 2035 — coinfanking · 2026-07-22
- Reddit user runs Nemotron Ultra 550B across aging MI50 and P40 GPU rigs — Old_Grapefruit8774 · 2026-07-22
- SK Hynix denies talks to buy Intel’s Ohio fab after market rumors — rwang07 · 2026-07-22
- Ineffable Labs takes delivery of a Vera Rubin NVL72 cluster — deanwball · 2026-07-22
- Chemical Giant Buys Its Way Into the GPU Rack Market — shashib · 2026-07-22
- GPT-6 is said to be near, with OpenAI betting on faster inference and custom chips — haider1 · 2026-07-22