SGLang Partners with Google and RadixArk to Deliver Native TPU Inference for LLMs

ying11231 · x · 2026-07-31

Open-source LLM inference framework SGLang announced a collaboration with Google Cloud and RadixArk to optimize inference for large language and diffusion models on TPUs.

Currently, the JAX-based SGL-JAX offers a production-level solution, delivering fast, native TPU inference for major model families like Gemma, Qwen, and DeepSeek. This partnership will enable TPUs to unlock SGLang's full production feature set, including 5D parallelism, Radix Cache, quantization, and speculative decoding. The team is also developing SGL-torchtpu to bring a PyTorch-native path to TPU inference.

Related event: SGLang Partners with Google Cloud to Support TPU Inference(3 posts)→

Original post →

More from Infra

Infra channel →