R1 runs at 150 tps, 7x faster than DeepSeek's serving

teortaxesTex · x · 2026-08-16

Discussion on inference performance: R1 runs at 150 tps on Rubin, about 7x faster than DeepSeek's official serving, reaching 60% of peak throughput and nearly 1 t/s/watt. Also mentions 200k is lower bound for pure inference; throttling Flash (≈GLM 5.2 tier) to 75 tps could support 200 agents per NPU, 1.6M per SuperPoD. Emphasizes perfect fine-grained replayability provides massive training signal, like an infinite data machine.

Related event: GLM's Rapid Updates and DeepSeek's Compute, Inference Speed Spark Debate(3 posts)→

Original post →

More from Infra

Infra channel →