SkewAdam Slashes MoE Optimizer Memory by 97.4%
A new preprint introduces SkewAdam, a tiered optimizer that reduces MoE training memory usage by 97.4%, compressing a 6.78B parameter model's optimizer state to just 1.29GB to fit on a 40GB GPU.
2026-07-22 ~ 2026-07-22 · 2 related posts
- SkewAdam cuts MoE optimizer memory by 97.4% and fits 6.78B on a 40GB GPU — Kooky-Ad-4124 · 2026-07-22
1 near-duplicate retellings: Nuemaan Malik