SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

ByteDance-Seed · hf · 2026-09-02

ByteDance presents SMELT, investigating how looping middle layers in sparse Mixture-of-Experts Transformers improves training efficiency and downstream performance. This method optimizes architecture while matching per-token FLOPs, parameters, and cache budgets.

Original post →

More from Research

Research channel →