SkewAdam Slashes MoE Optimizer Memory by 97.4%

A new preprint introduces SkewAdam, a tiered optimizer that reduces MoE training memory usage by 97.4%, compressing a 6.78B parameter model's optimizer state to just 1.29GB to fit on a 40GB GPU.

2026-07-22 ~ 2026-07-22 · 2 related posts

1 near-duplicate retellings: Nuemaan Malik