TPUv7 Hits 700 tok/s in vLLM Test, Beating GB200 by 56%

vLLM maintainers measured TPUv7 running Kimi K3 at 700 tok/s per user with Megakernel optimizations, 56% faster than Nvidia's GB200 NVL72, while inferact open-sourced the first TPU inference megakernel reaching 709 tokens/s.

2026-09-24 ~ 2026-09-24 · 2 related posts