Prime Intellect cuts GLM-5.2 RL weight transfer from 86s to 4s with NIXL and ModelExpress
samsja19 · x · 2026-09-05
Prime Intellect rebuilt weight synchronization for trillion-parameter RL training using NIXL and ModelExpress, cutting GLM-5.2 weight transfer from 86 seconds to 4 seconds. After prime-rl got sub-5-minute step times on 28 H200 nodes, the multi-TB trainer-to-sampler weight sync (stuck at 60-90s) became the bottleneck. The new approach also removes static process-group constraints, enabling fault-tolerant, elastic inference, compared against NCCL and filesystem/delta-based approaches and their tradeoffs.
More from Infra
- AMD, Cisco and Saudi Arabia's HUMAIN deploy MI335X GPUs, planning up to 250MW — Beth_Kindig · 2026-09-05
- Reef launches inference-native infra that serves self-improving agents without downtime — pliang279 · 2026-09-05
- From 1D to 2D int8 kernels: a hands-on GPU internals learning path — goyal__pramod · 2026-09-05
- Google DeepMind Publishes Free Book on Scaling LLMs Across TPUs and GPUs — goyal__pramod · 2026-09-05
- Prime Super Flash MoE: 1.2x BF16 and 1.6x MXFP8 speedups over upstream on B200 — retr0jirachi · 2026-09-05
- How to Estimate tokens/sec on Your Hardware: The VRAM Bandwidth Formula — Pyrolistical · 2026-09-05