PyTorch unveils TorchTPU backend so vLLM and SGLang run natively on TPUs

PyTorch · x · 2026-10-02

PyTorch announced vLLM and SGLang engines built on a new PyTorch-native TPU backend, TorchTPU. At PyTorch Conference North America in San Jose, Qi Zhou (Google), Colin Taylor and Angela Yi (Meta) will explain how they made the existing serving stacks work natively on TPUs while preserving schedulers, batching, OpenAI-compatible APIs, and torch.compile workflows familiar to GPU users.

Related event: PyTorch Unveils TorchTPU, Enabling Native vLLM and SGLang on TPU(2 posts)→

Original post →

More from Infra

Infra channel →