SGLang Comes to Google TPUs: Enabling Seamless Migration for Major LLMs
BanghuaZ · x · 2026-07-31
Google Cloud and RadixArk have partnered to bring the open-source inference framework SGLang to TPUs.
- Current Support: Developers can run SGLang on the latest TPU generations via SGL-JAX, supporting major LLMs like Gemma, Qwen, DeepSeek, Kimi, and Grok, alongside diffusion models like Wan and Flux.
- Future Roadmap: Later this year, RadixArk will roll out SGL-torchtpu, a PyTorch-native backend featuring full parallelism capabilities (data, tensor, expert, context, pipeline) for full-size frontier models on multi-host TPUs.
- Ecosystem Goal: Pledges Day 0 support for new open models on TPU, making TPUs a drop-in, cost-efficient path for frontier inference.
Related event: SGLang Partners with Google Cloud to Support TPU Inference(3 posts)→
More from Infra
- Running Inkling Small on a Single Node: Native Voice Interaction Under 500ms — andimarafioti · 2026-07-31
- LLM Inference Costs Plunge: Token Prices Drop to 1/13th in Four Months — charliermarsh · 2026-07-31
- Samsung Earnings: Agentic AI Drives Surge in Enterprise SSD Demand — davidyin44 · 2026-07-31
- vLLM Releases Inkling-Small Deployment Guide: Runs on Minimum 180GB VRAM — vllm_project · 2026-07-31
- Malaysian Activists Successfully Halt Data Centre Construction — im_mansigupta · 2026-07-31
- Meta's AI Infrastructure Lease Obligations Surge 53% in Three Months to Nearly $279 Billion — Polymarket · 2026-07-31