PyTorch ships TorchTPU: vLLM and SGLang now run natively on Google TPUs
PyTorch · x · 2026-10-02
PyTorch announced TorchTPU, a PyTorch-native TPU backend that lets the vLLM and SGLang serving engines run natively on TPUs while preserving schedulers, batching, OpenAI-compatible APIs, and torch.compile workflows. Google and Meta engineers will share production lessons — compile times, device placement under tracing, KV cache layout, cross-chip collectives — at PyTorch Conference North America in San Jose.
Related event: PyTorch Unveils TorchTPU, Enabling Native vLLM and SGLang on TPU(2 posts)→
More from Infra
- Benchmark Backs Chip Startup Tendrils Compute in Round That Could Top $1B Valuation — Sethwinterroth · 2026-10-02
- Modal Runtime ships multi-node clusters, VM sandboxes and sticky sessions — graceisford · 2026-10-02
- How to run local T2V/I2V/V2V video AI on an RTX 5070 with only 12GB VRAM — AirSea2062 · 2026-10-02
- Lightmatter CEO explains how spare lasers and optical switching route around failures in AI systems — BenBajarin · 2026-10-02
- Nebius Outperforms CoreWeave by 8X Despite 5X Less Revenue: What's Going On — Beth_Kindig · 2026-10-02
- Pedro Domingos: data centers lower electricity rates by paying the same overhead as residents — pmddomingos · 2026-10-02