Cross-Datacenter RL: Modal Shrinks 500GB Weight Syncs to 500MB

AI Engineer · youtube · 2026-08-11

Nan Jiang from Modal detailed an engineering optimization for cross-datacenter Reinforcement Learning (RL). In RL training, shipping a 500GB checkpoint to rollout fleets in other regions takes minutes to hours, stalling weight updates.

Core Finding: Fewer than 1% of served weights actually change between consecutive versions. This isn't due to gradient sparsity (99% get non-zero gradients), but because of the gap between the precision floor of low-precision serving and the tiny Adam step size, an effect called Adam absorption.

Implementation:

Original post →

More from Infra

Infra channel →