NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer
dair_ai · x · 2026-10-09
An NVIDIA paper found that only 0.6%–1.2% of model weights change per RL training step, yet standard practice copies the full checkpoint to the rollout cluster after every update — 87.5 minutes for a 1T model across two AWS regions.
How NeMo-DCR works:
- Sends only the changed weights; the rollout cluster receives bit-identical results to a full copy
- Maps changes from training shards directly into checkpoint layout, encoding them as XOR masks or overwrites
- Streams deltas through a relay tree while still being built, with joint commits and retries to handle mid-transfer failures
Results: a 1T refit at 3% change rate takes 150 seconds instead of 87.5 minutes; across 30B–1T models, it is 12–40x faster than full-checkpoint transfer.
Related event: NVIDIA NeMo-DCR Speeds Up Trillion-Parameter RL Weight Sync by Up to 40x(2 posts)→
More from Infra
- Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory — _akhaliq · 2026-10-09
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- Zyphra Speeds MoE Expert Routing Communication 2.63x on AMD MI300X GPUs — QuentinAnthon15 · 2026-10-09