SMELT: Looped MoE Transformers Outperform Under Compute-Matched Budgets
scaling01 · x · 2026-09-02
A new arXiv paper, "SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers," proposes an architecture that loops the middle half of layers twice. While matching baseline FLOPs, non-embedding parameters, and KV cache, SMELT achieves faster loss reduction, saving 6.8--18.0% of training FLOPs. Gains are most significant on Code benchmarks and scale with context length.
Related event: ByteDance's SMELT: Looping MoE Layers Cuts Training Compute Up to 18%(4 posts)→
More from Research
- Offline Rubric Synthesis Plus Refinement Loops: A Practical Reward Hacking Mitigation — stochasticchasm · 2026-09-22
- Frontend design framed as visual agent task with groupwise relative grading — stochasticchasm · 2026-09-22
- Team reportedly plans to open source 7,000 RL training environments — airesearch12 · 2026-09-22
- Why RL generalizes to reasoning but not literary writing, per AI researchers — phl43 · 2026-09-22
- Rethinking Policy Gradients: Score Centering Skips Importance Sampling Entirely — brandondamos · 2026-09-22
- Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition — hunarbatra · 2026-09-22