PyTorchCon talk: MoRI + vLLM brings RDMA KV-cache transfer and wide expert parallelism to AMD
PyTorch · x · 2026-10-11
At PyTorch Conference North America 2026, three AMD engineers will present "MoRI + vLLM: Wide Expert Parallelism and RDMA KV-Cache Transfer for Disaggregated MoE Serving on AMD."
- Running large MoE models across pods raises two challenges: expert dispatch/combine traffic, and moving the KV cache between prefill and decode.
- The work brings the MoRI communication stack into vLLM, enabling Wide Expert Parallelism across pods via RDMA, plus KV-cache transfer for disaggregated prefill/decode on AMD Instinct GPUs.
- They will also cover correctness and stability issues when spanning pods, with vLLM benchmarks on long, high-concurrency workloads.
More from Infra
- DeepSeek v4.1 reportedly boosts long-context prefill throughput by ~40% — HankYeomans · 2026-10-11
- GPU Tsunami: how advanced packaging is reshaping the semiconductor test market — BenBajarin · 2026-10-11
- Optical testing is the key bottleneck for scaling co-packaged optics deployment — BenBajarin · 2026-10-11
- Cboe's former HQ to be converted into a 33MW data center with above-market rents — deanwball · 2026-10-11
- Microsoft open-sources bitnet.cpp, running 100B models on CPU at 6.17x speed — HildeKuehne · 2026-10-11
- Qualcomm CEO: AI firms want phones running 100B-parameter models continuously by 2028 — zephyr_z9 · 2026-10-11