ByteDance's SMELT: Looping MoE Layers Cuts Training Compute Up to 18%
ByteDance Seed, Tsinghua and TokenWave introduced SMELT, a recurrent sparse MoE Transformer that loops middle layers twice, cutting training FLOPs by 6.8%-18% under compute-matched scaling laws while improving downstream performance.
2026-09-02 ~ 2026-09-03 · 4 related posts
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers — ByteDance-Seed · 2026-09-02
- SMELT paper: looping middle layers twice in MoE saves 6.8-18% training FLOPs — iScienceLuvr · 2026-09-02
- Looped Transformer SMELT cuts training FLOPs 6.8-18%, tipped for GPT-6 — mark_k · 2026-09-03
1 near-duplicate retellings: scaling01