SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
ByteDance-Seed · hf · 2026-09-02
ByteDance presents SMELT, investigating how looping middle layers in sparse Mixture-of-Experts Transformers improves training efficiency and downstream performance. This method optimizes architecture while matching per-token FLOPs, parameters, and cache budgets.
More from Research
- YC Paper Club Call: Optical Compute, Diamond Chips, and Bio-GPUs — ycombinator · 2026-09-02
- Recurrent Activations Raise AI Monitoring Challenges — RyanGreenblatt · 2026-09-02
- CrossFeat: Bridging Imaging Modalities in Feature Space — zhenjun_zhao · 2026-09-02
- Gravity-Prior Driven Decoupling for Robust Pose Estimation — zhenjun_zhao · 2026-09-02
- DualDiff3D: Dual Diffusion Priors for Robust 3DGS — zhenjun_zhao · 2026-09-02
- Meta paper: Agents complete tasks but fail to prevent catastrophic actions like factory resets — rohanpaul_ai · 2026-09-02