SkewAdam cuts MoE optimizer memory by 97.4% and fits 6.78B on a 40GB GPU

Kooky-Ad-4124 · reddit · 2026-07-22

A new preprint introduces SkewAdam, a tiered optimizer designed to slash the memory cost of Mixture-of-Experts training.

Original post →

More from Infra

Infra channel →