vLLM Sharded Weight Transfer Hits 7.53s for 1T Params Model

TheZachMueller · x · 2026-08-26

The vLLM team, in collaboration with SkyRL, has implemented a native sharded weight transfer engine using Ray Direct Transfer (RDT) and NIXL. This addresses bottlenecks in syncing weights for large-scale online RL. The implementation supports dense, MoE (fused or per-expert), and quantized models. On a cluster of 48 8xH100 nodes, sharded weight transfer for the Kimi K2 model (1T params, BF16) takes just 7.53 seconds. The work also features optimizations overlapping preprocessing with transport and a fault-tolerant rollout demonstration.

Original post →

More from Infra

Infra channel →