Dynamic REAP swaps 4 experts every 256 tokens, keeping only 25% in GPU memory

Tim_Dettmers · x · 2026-09-22

Compressed models deploy with Dynamic REAP: even keeping only 25% of experts per layer, routing probabilities of removed experts faithfully indicate which experts the current distribution needs, guiding async transfers from CPU RAM/NVMe to GPU. Four experts are swapped every 256 tokens at token boundaries without disturbing inference. Surprisingly, this yields better out-of-distribution generalization than fitting REAP on the test set — MoE routing distributions change frequently as documents are processed (agent work alternates docs, code, user interaction).

Related event: Dettmers Explains Iterative Sensitivity Probing and Dynamic REAP(2 posts)→

Original post →

More from Infra

Infra channel →