Prime Intellect cuts GLM-5.2 RL weight transfer from 86s to 4s with NIXL and ModelExpress

samsja19 · x · 2026-09-05

Prime Intellect rebuilt weight synchronization for trillion-parameter RL training using NIXL and ModelExpress, cutting GLM-5.2 weight transfer from 86 seconds to 4 seconds. After prime-rl got sub-5-minute step times on 28 H200 nodes, the multi-TB trainer-to-sampler weight sync (stuck at 60-90s) became the bottleneck. The new approach also removes static process-group constraints, enabling fault-tolerant, elastic inference, compared against NCCL and filesystem/delta-based approaches and their tradeoffs.

Related event: Prime Intellect's RL stack adopts NIXL, cutting 800B-model weight transfer 9x from 86s to under 4s(2 posts)→

Original post →

More from Infra

Infra channel →