SGLang Partners with Google and RadixArk to Deliver Native TPU Inference for LLMs
ying11231 · x · 2026-07-31
Open-source LLM inference framework SGLang announced a collaboration with Google Cloud and RadixArk to optimize inference for large language and diffusion models on TPUs.
Currently, the JAX-based SGL-JAX offers a production-level solution, delivering fast, native TPU inference for major model families like Gemma, Qwen, and DeepSeek. This partnership will enable TPUs to unlock SGLang's full production feature set, including 5D parallelism, Radix Cache, quantization, and speculative decoding. The team is also developing SGL-torchtpu to bring a PyTorch-native path to TPU inference.
Related event: SGLang Partners with Google Cloud to Support TPU Inference(3 posts)→
More from Infra
- Meta's AI Infrastructure Lease Obligations Surge 53% in Three Months to Nearly $279 Billion — Polymarket · 2026-07-31
- SGLang Comes to Google TPUs: Enabling Seamless Migration for Major LLMs — BanghuaZ · 2026-07-31
- Running GLM-5.2 and Kimi K3 Locally on RTX 5090: Hits 129 tok/s — markjeffrey · 2026-07-31
- AI Inference Demand Shatters Forecasts: Google Token Consumption Up 330X in Two Years — Beth_Kindig · 2026-07-31
- Investor: The Current AI Boom Would Not Have Been Possible Without Crypto — davidyin44 · 2026-07-31
- OpenAI Inference Compute Estimated at 2-3GW, Far Exceeding Chinese Open Source — zephyr_z9 · 2026-07-31