Open-sourced TPU megakernel runs Kimi K3 at 709 tokens/s, beating GB200

vllm_project · x · 2026-09-24

The inferact team open-sourced what they believe is the first TPU inference megakernel, reaching 709 tokens/s on low-concurrency decode for Kimi K3, versus 450 tokens/s on their GB200 baseline (both with DSpark speculative decoding). All 92 of K3's MoE layers run in a single Pallas kernel, with weight prefetching across layer boundaries so transfers for one layer overlap with computation in the previous one. Without speculative decoding it is roughly 1.4-2x the GB200 baseline at batch sizes 1 through 8. The vLLM project amplified the release.

Original post →

More from Infra

Infra channel →