vLLM Team's Inferact Runs Kimi K3 57% Faster on 16 TPUs Than GB200 With Open-Sourced Megakernel
量子位 · wechat · 2026-09-26
Inferact, founded by the original vLLM team ($150M seed at $800M valuation), open-sourced a megakernel that runs Kimi K3 at 709 tokens/s on 16 TPUv7 — 57% faster than 16 GB200 at 452 tokens/s, with zero accuracy loss. The trick: fusing all 92 MoE layers into one Pallas program, eliminating kernel boundaries and enabling cross-layer weight prefetch. TPU's software-managed 64MiB VMEM beats GB200's hardware-scheduled 38MiB despite lower HBM bandwidth, showing the gap is software, not silicon.
More from Infra
- Chart claims $6.5 trillion in current and future liabilities between LLM labs and hyperscalers — AIFlow_ML · 2026-09-26
- 128GB local LLM server forges 80M tokens a month on just $6 of electricity — No-Fuel-9202 · 2026-09-26
- Community poll of 584 ballots crowns oMLX (57%) the top MLX engine on Apple Silicon, ahead of MLX-Serve and LM Studio — andrejusb · 2026-09-26
- Big 4 AI capex hits 2.4% of GDP, twice the peak of the telecom boom 25 years ago — deedydas · 2026-09-26
- Vitalik pitches local Qwen AI with 100+ GB of offline data to de-centralize Ethereum nodes and IPFS — DavideCrapis · 2026-09-26
- If Jensen's math holds — 1GW = $100B — the AI compute bill looks like a problem — AIFlow_ML · 2026-09-26