vLLM Sharded Weight Transfer Hits 7.53s for 1T Params Model
TheZachMueller · x · 2026-08-26
The vLLM team, in collaboration with SkyRL, has implemented a native sharded weight transfer engine using Ray Direct Transfer (RDT) and NIXL. This addresses bottlenecks in syncing weights for large-scale online RL. The implementation supports dense, MoE (fused or per-expert), and quantized models. On a cluster of 48 8xH100 nodes, sharded weight transfer for the Kimi K2 model (1T params, BF16) takes just 7.53 seconds. The work also features optimizations overlapping preprocessing with transport and a fault-tolerant rollout demonstration.
More from Infra
- OpenAI's Jalapeño Chip Reportedly Beats Nvidia Blackwell on Perf/Watt — petrusenko_max · 2026-08-26
- Why AI projects fail in 2026? Shift to infra, talent, and ROI — mikeflache · 2026-08-26
- New API pricing drops to $1 per 1K requests, built-in search and fetch tools — testingcatalog · 2026-08-26
- Apple M5 Ultra cluster hits 4.8TB/s bandwidth, rivaling data center GPUs — awnihannun · 2026-08-26
- Snowflake reveals Agent hidden costs, reducing trial cost by 33%-45% — StasBekman · 2026-08-26
- M5 Ultra Studio pricing is wonderfully broken: 2-3x DGX Spark value for local AI — SumitGup · 2026-08-26