vLLM Team's Inferact Runs Kimi K3 57% Faster on 16 TPUs Than GB200 With Open-Sourced Megakernel

量子位 · wechat · 2026-09-26

Inferact, founded by the original vLLM team ($150M seed at $800M valuation), open-sourced a megakernel that runs Kimi K3 at 709 tokens/s on 16 TPUv7 — 57% faster than 16 GB200 at 452 tokens/s, with zero accuracy loss. The trick: fusing all 92 MoE layers into one Pallas program, eliminating kernel boundaries and enabling cross-layer weight prefetch. TPU's software-managed 64MiB VMEM beats GB200's hardware-scheduled 38MiB despite lower HBM bandwidth, showing the gap is software, not silicon.

Original post →

More from Infra

Infra channel →