Dynamic REAP swaps 4 experts every 256 tokens, keeping only 25% in GPU memory
Tim_Dettmers · x · 2026-09-22
Compressed models deploy with Dynamic REAP: even keeping only 25% of experts per layer, routing probabilities of removed experts faithfully indicate which experts the current distribution needs, guiding async transfers from CPU RAM/NVMe to GPU. Four experts are swapped every 256 tokens at token boundaries without disturbing inference. Surprisingly, this yields better out-of-distribution generalization than fitting REAP on the test set — MoE routing distributions change frequently as documents are processed (agent work alternates docs, code, user interaction).
Related event: Dettmers Explains Iterative Sensitivity Probing and Dynamic REAP(2 posts)→
More from Infra
- H Company trains computer-use agents on SkyPilot: thousands of sub-second sandboxes — skypilot_org · 2026-09-23
- MiniMax H3 video gen runs locally on M5 Ultra: 768p in ~2m22s with optimizations — bakawolf123 · 2026-09-22
- Cisco: Agentic AI to Drive 9X Enterprise Traffic Growth by 2035 vs 2.5X Without — Beth_Kindig · 2026-09-22
- Cloudflare ships Vary support in Cache Rules to tame HTTP's 'ugliest' header — threepointone · 2026-09-22
- Fluidstack breaks ground on $4B Texas data center campus planned for up to 1.5 GW — MxMnr · 2026-09-22
- Right-size GenAI endpoints: SageMaker concurrency sweeps with Nemotron-3 Nano 30B — AWS ML Blog · 2026-09-22