Google Cloud and Inferact partner to make TPU a first-class vLLM target
vllm_project · x · 2026-09-15
- Google Cloud and Inferact (the company behind vLLM) announced an engineering partnership to make TPU a first-class citizen in vLLM, with all work open-sourced.
- Scope: production serving features and optimized kernels, a native PyTorch path via TorchTPU, and day-0 support for frontier model releases.
- A community program offers shared TPU capacity for open-source contributors plus dedicated review/design help from Inferact's core vLLM maintainers.
- Context: vLLM supports 5,000+ model architectures with 3,000+ contributors; Anthropic has committed to scale to as many as 1 million TPUs, and the latest Ironwood generation is built for inference.
More from Infra
- OpenAI engineers: kernel optimization cut GPT-5.6 Sol serving cost by 20% — TheTuringPost · 2026-09-15
- Devin left alone with Modal H100s cuts training kernel peak memory 46% and latency 52% — AAAzzam · 2026-09-15
- Why Macs quietly win at local AI: unified memory beats RTX 5090 and accessibility APIs power better computer use — dotey · 2026-09-15
- 5 production apps, 2M monthly requests for $6: why developers are going all-in on Cloudflare — viksit · 2026-09-15
- Mixing a 3090 with an Intel Arc B70 for 56GB VRAM local LLM inference? — overand · 2026-09-15
- How Abnormal AI Screens Billions of Emails with Bedrock AgentCore Sandboxes — AWS ML Blog · 2026-09-15