vLLM Tests Show TPUv7 Hits 700 tok/s/user, 56% Faster Than NVIDIA GB200 NVL72
woosuk_k · x · 2026-09-24
Per SemiAnalysis, vLLM maintainers demonstrated that TPUv7 can reach 700 tok/s per user on Kimi K3 via Megakernel optimization — 56% better throughput than NVIDIA's GB200 NVL72. SemiAnalysis argues that TPU software externalization is now moving at full speed, making this a key inflection in the accelerator landscape to watch.
Related event: TPUv7 Hits 700 tok/s in vLLM Test, Beating GB200 by 56%(2 posts)→
More from Infra
- Nvidia CEO Jensen Huang expects to sell twice as many chips next year — Beth_Kindig · 2026-09-24
- Running Qwen locally with Hermes agent in 10 minutes on an old gaming laptop — markjeffrey · 2026-09-24
- Chris Lattner makes first Snapdragon Summit appearance as Qualcomm EVP after Modular deal — samcharrington · 2026-09-24
- MongoDB 3.0: a decade of sales-led growth gets rebuilt — thedealdirector · 2026-09-24
- rauchg: every successful agent needs brain, hands and files — decouple them in the cloud — cramforce · 2026-09-24
- Meta's $12B Compute Commitment Gives Nebius Upside on Up to $15B More — AccBalanced · 2026-09-24