vLLM adds Elastic Expert Parallelism: grow/shrink MoE GPU pools under live traffic

PyTorch · x · 2026-09-26

PyTorch announced Elastic Expert Parallelism in vLLM, which lets operators add or remove GPUs from an active Mixture-of-Experts deployment during live traffic with minimal serving interruption and downtime.

NVIDIA's Itay Alroy will present 'Elastic Expert Parallelism in vLLM' at PyTorch Conference North America 2026, covering the architecture, key implementation details, open challenges, and roadmap—including what happens when EP size changes and how NIXL EP enables grow/shrink under live traffic.

Original post →

More from Infra

Infra channel →