PyTorch unveils TorchTPU backend so vLLM and SGLang run natively on TPUs
PyTorch · x · 2026-10-02
PyTorch announced vLLM and SGLang engines built on a new PyTorch-native TPU backend, TorchTPU. At PyTorch Conference North America in San Jose, Qi Zhou (Google), Colin Taylor and Angela Yi (Meta) will explain how they made the existing serving stacks work natively on TPUs while preserving schedulers, batching, OpenAI-compatible APIs, and torch.compile workflows familiar to GPU users.
- Engines stay upstream, so new models and features run on TPU without separate TPU-specific implementations
- Production lessons to be shared: compile time, device placement under tracing, KV cache layout, and cross-chip collectives
Related event: PyTorch Unveils TorchTPU, Enabling Native vLLM and SGLang on TPU(2 posts)→
More from Infra
- Benchmark Backs Chip Startup Tendrils Compute in Round That Could Top $1B Valuation — Sethwinterroth · 2026-10-02
- Modal Runtime ships multi-node clusters, VM sandboxes and sticky sessions — graceisford · 2026-10-02
- How to run local T2V/I2V/V2V video AI on an RTX 5070 with only 12GB VRAM — AirSea2062 · 2026-10-02
- Nebius Outperforms CoreWeave by 8X Despite 5X Less Revenue: What's Going On — Beth_Kindig · 2026-10-02
- Pedro Domingos: data centers lower electricity rates by paying the same overhead as residents — pmddomingos · 2026-10-02
- PyTorch ships TorchTPU: vLLM and SGLang now run natively on Google TPUs — PyTorch · 2026-10-02