TPUv7 Hits 700 tok/s in vLLM Test, Beating GB200 by 56%
vLLM maintainers measured TPUv7 running Kimi K3 at 700 tok/s per user with Megakernel optimizations, 56% faster than Nvidia's GB200 NVL72, while inferact open-sourced the first TPU inference megakernel reaching 709 tokens/s.
2026-09-24 ~ 2026-09-24 · 2 related posts
- Open-sourced TPU megakernel runs Kimi K3 at 709 tokens/s, beating GB200 — vllm_project · 2026-09-24
- vLLM Tests Show TPUv7 Hits 700 tok/s/user, 56% Faster Than NVIDIA GB200 NVL72 — woosuk_k · 2026-09-24