xAI SGLang Lead: Deploying Grok Across 100,000 GPUs
FinanceYF5 · x · 2026-07-07
In a 23-minute video, a former Berkeley PhD and xAI's SGLang lead detailed the engineering architecture for deploying Grok on a 100,000-GPU scale. Key optimizations include splitting Prefill and Decode phases, sharding MoE experts across GPUs, routing tokens by expert to cut communication overhead, and overlapping communication with computation to hide latency. This enables service pricing lower than the DeepSeek API. The breakdown of large-scale MoE inference trade-offs serves as a deep dive into state-of-the-art inference stacks.
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11