ByteDance's SMELT: Looping MoE Layers Cuts Training Compute Up to 18%

ByteDance Seed, Tsinghua and TokenWave introduced SMELT, a recurrent sparse MoE Transformer that loops middle layers twice, cutting training FLOPs by 6.8%-18% under compute-matched scaling laws while improving downstream performance.

2026-09-02 ~ 2026-09-03 · 4 related posts

1 near-duplicate retellings: scaling01