SMELT: Looped MoE Transformers Outperform Under Compute-Matched Budgets

scaling01 · x · 2026-09-02

A new arXiv paper, "SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers," proposes an architecture that loops the middle half of layers twice. While matching baseline FLOPs, non-embedding parameters, and KV cache, SMELT achieves faster loss reduction, saving 6.8--18.0% of training FLOPs. Gains are most significant on Code benchmarks and scale with context length.

Related event: ByteDance's SMELT: Looping MoE Layers Cuts Training Compute Up to 18%(4 posts)→

Original post →

More from Research

Research channel →