xAI SGLang Lead: Deploying Grok Across 100,000 GPUs

FinanceYF5 · x · 2026-07-07

In a 23-minute video, a former Berkeley PhD and xAI's SGLang lead detailed the engineering architecture for deploying Grok on a 100,000-GPU scale. Key optimizations include splitting Prefill and Decode phases, sharding MoE experts across GPUs, routing tokens by expert to cut communication overhead, and overlapping communication with computation to hide latency. This enables service pricing lower than the DeepSeek API. The breakdown of large-scale MoE inference trade-offs serves as a deep dive into state-of-the-art inference stacks.

Original post →

More from Infra

Infra channel →