vLLM thread (4/5): locality-domain MoE sharding speeds up decode 1.2x
vllm_project · x · 2026-10-10
Part 4 of vLLM's thread explains a concrete optimization: since Ampere, NVIDIA GPUs have non-uniform global memory access, exposed via locality domains in CUDA 13.4 — HBM is partitioned so SMs read their local partition fastest. vLLM shards MoE FC1/FC2 weights column-wise and uses Green Contexts to launch one kernel per domain, cutting MiniMax M3 MoE decode latency by up to 1.2x.
More from Infra
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- DIY hybrid GPU/CPU/SSD rig cuts DeepSeek TTFT from 75s to 8.9s at 16K prefill — HankYeomans · 2026-10-10
- vLLM lands NVIDIA Vera Rubin support, hitting 7.8x GB200 throughput on MiniMax M3 — vllm_project · 2026-10-10