PyTorchCon talk: MoRI + vLLM brings RDMA KV-cache transfer and wide expert parallelism to AMD

PyTorch · x · 2026-10-11

At PyTorch Conference North America 2026, three AMD engineers will present "MoRI + vLLM: Wide Expert Parallelism and RDMA KV-Cache Transfer for Disaggregated MoE Serving on AMD."

Original post →

More from Infra

Infra channel →