PyTorch ships TorchTPU: vLLM and SGLang now run natively on Google TPUs

PyTorch · x · 2026-10-02

PyTorch announced TorchTPU, a PyTorch-native TPU backend that lets the vLLM and SGLang serving engines run natively on TPUs while preserving schedulers, batching, OpenAI-compatible APIs, and torch.compile workflows. Google and Meta engineers will share production lessons — compile times, device placement under tracing, KV cache layout, cross-chip collectives — at PyTorch Conference North America in San Jose.

Related event: PyTorch Unveils TorchTPU, Enabling Native vLLM and SGLang on TPU(2 posts)→

Original post →

More from Infra

Infra channel →