vLLM thread (4/5): locality-domain MoE sharding speeds up decode 1.2x

vllm_project · x · 2026-10-10

Part 4 of vLLM's thread explains a concrete optimization: since Ampere, NVIDIA GPUs have non-uniform global memory access, exposed via locality domains in CUDA 13.4 — HBM is partitioned so SMs read their local partition fastest. vLLM shards MoE FC1/FC2 weights column-wise and uses Green Contexts to launch one kernel per domain, cutting MiniMax M3 MoE decode latency by up to 1.2x.

Original post →

More from Infra

Infra channel →