Open-sourced TPU megakernel runs Kimi K3 at 709 tokens/s, beating GB200
vllm_project · x · 2026-09-24
The inferact team open-sourced what they believe is the first TPU inference megakernel, reaching 709 tokens/s on low-concurrency decode for Kimi K3, versus 450 tokens/s on their GB200 baseline (both with DSpark speculative decoding). All 92 of K3's MoE layers run in a single Pallas kernel, with weight prefetching across layer boundaries so transfers for one layer overlap with computation in the previous one. Without speculative decoding it is roughly 1.4-2x the GB200 baseline at batch sizes 1 through 8. The vLLM project amplified the release.
More from Infra
- Prime Intellect releases Prime Sandboxes, MicroVMs purpose-built for RL training — TheZachMueller · 2026-09-24
- Qualcomm and Liquid AI CEOs discuss co-designing hardware and models for on-device AI — samcharrington · 2026-09-24
- Zilliz CTO: agents make the enterprise data layer impossible to ignore — No_Engineer_1224 · 2026-09-24
- Running Android emulator + Chrome with 60fps streaming in a $0.072/hr cloud VM for always-on agents — cem2ran · 2026-09-24
- NVIDIA's DGX Spark Appears Unavailable, May Never Return to Sale — GabGarrett · 2026-09-24
- Not every AI task needs an LLM: 'decide' may become a standard model call — bigdata · 2026-09-24