DeepSeek post pairs inference-system optimizations with a 545% theoretical margin analysis

teortaxesTex · x · 2026-07-22

A repost about DeepSeek combines an inference-system overview with a separate analysis of its theoretical revenue and expenses. The screenshot cites throughput optimizations such as cross-node EP batch scaling, computation-communication overlap, and load balancing, plus claims about H800 token throughput and a theoretical 545% cost profit margin.

The attached chart and paper page also compare network and latency choices such as RoCE vs. InfiniBand, discussing congestion control, routing, and GPU-direct communication tradeoffs for large-scale AI workloads.

Original post →

More from Infra

Infra channel →