R1 runs at 150 tps, 7x faster than DeepSeek's serving
teortaxesTex · x · 2026-08-16
Discussion on inference performance: R1 runs at 150 tps on Rubin, about 7x faster than DeepSeek's official serving, reaching 60% of peak throughput and nearly 1 t/s/watt. Also mentions 200k is lower bound for pure inference; throttling Flash (≈GLM 5.2 tier) to 75 tps could support 200 agents per NPU, 1.6M per SuperPoD. Emphasizes perfect fine-grained replayability provides massive training signal, like an infinite data machine.
Related event: GLM's Rapid Updates and DeepSeek's Compute, Inference Speed Spark Debate(3 posts)→
More from Infra
- llama.cpp integrates Dots3 Note model, scoring 78.4 on SWE-bench Verified — victormustar · 2026-08-16
- Weaviate adds test-time compute scaling to Search Mode, boosting retrieval performance significantly — dl_weekly · 2026-08-16
- New Book: Algorithms for Modern Hardware Open-Sourced on GitHub — thehiphopswami · 2026-08-16
- Comfy Kitchen Attention Speeds Up MiniMax H3 on AMD GPUs — God_Hand_9764 · 2026-08-16
- Quality difference between Q8_0 and UD-Q6_K_XL quantization — AnimalPuzzleheaded71 · 2026-08-16
- Gavin Baker: NVIDIA is becoming the "central bank of AI" — VibeMarketer_ · 2026-08-16