vLLM Tests Show TPUv7 Hits 700 tok/s/user, 56% Faster Than NVIDIA GB200 NVL72

woosuk_k · x · 2026-09-24

Per SemiAnalysis, vLLM maintainers demonstrated that TPUv7 can reach 700 tok/s per user on Kimi K3 via Megakernel optimization — 56% better throughput than NVIDIA's GB200 NVL72. SemiAnalysis argues that TPU software externalization is now moving at full speed, making this a key inflection in the accelerator landscape to watch.

Related event: TPUv7 Hits 700 tok/s in vLLM Test, Beating GB200 by 56%(2 posts)→

Original post →

More from Infra

Infra channel →