vLLM adds Elastic Expert Parallelism: grow/shrink MoE GPU pools under live traffic
PyTorch · x · 2026-09-26
PyTorch announced Elastic Expert Parallelism in vLLM, which lets operators add or remove GPUs from an active Mixture-of-Experts deployment during live traffic with minimal serving interruption and downtime.
NVIDIA's Itay Alroy will present 'Elastic Expert Parallelism in vLLM' at PyTorch Conference North America 2026, covering the architecture, key implementation details, open challenges, and roadmap—including what happens when EP size changes and how NIXL EP enables grow/shrink under live traffic.
More from Infra
- Glamsterdam will replace sync healing with state diffs, further speeding up Ethereum nodes — banteg · 2026-09-26
- Unverified claim: xAI's 200k GB300 cluster at 10% MFU sparks community pushback — teortaxesTex · 2026-09-26
- NVIDIA at $5.4 trillion is now worth more than the entire UK or French stock market — iamfakhrealam · 2026-09-26
- DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests — teortaxesTex · 2026-09-26
- A Curated Paper List for Learning Distributed LLM Training and Inference — East-Muffin-6472 · 2026-09-26
- AI data center investors now favor real infrastructure over PowerPoint promises — TansuYegen · 2026-09-26