SGLang Partners with Google Cloud and RadixArk to Make TPU a Cost-Efficient Path for Frontier Inference

ying11231 · x · 2026-08-01

RadixArk, Google Cloud, and the SGLang community are collaborating to make TPU a drop-in, cost-efficient path for frontier LLM inference.

SGL-JAX already supports running major open-source models (like Gemma, Qwen, DeepSeek, etc.) on the latest TPU generations. This partnership brings SGLang's production-ready features—such as 5D parallelism, Radix Cache, and speculative decoding—to TPU. Furthermore, SGL-torchtpu will be introduced later this year as a PyTorch-native backend, enabling developers to run LLMs on TPUs using standard PyTorch tooling.

Related event: SGLang Partners with Google Cloud and RadixArk for TPU Inference(6 posts)→

Original post →

More from Infra

Infra channel →