Foil Flattens Looped MoE Experts to Cut Pretraining Loss by 0.012 nat at Equal Compute
Shouren Wang · hf · 2026-10-06
- Problem: Looped transformers reuse one layer block multiple times to squeeze more out of fixed parameters; sparse MoE activates a few of many experts. Foil bridges the two: how to loop a MoE.
- Method: With expert params and per-token expert compute fixed, Foil (1) flattens the experts—halving expert layers, doubling experts per layer and doubling passes so each routing decision picks from a larger pool—and (2) unties attention, giving each pass its own attention weights while experts/routers stay shared.
- Results: At 20B tokens every Foil model beats the unflattened baseline on pretraining loss; at 100B tokens loss improves monotonically with flattening, with the most flattened Foil ending 0.012 nat below baseline at equal params and compute, and downstream accuracy on par or better. Untied attention yields more balanced, more confident routing.
- Design guidance: Looping and widening expert layers amplify each other; routing confidence tracks healthy expert use better than load balance; sparse looped MoEs should use more experts per layer and more passes.
- Code and configs open-sourced at github.com/SR-A-W/how-to-loop-moe.
More from Research
- Scholars Push Back on Reported arXiv Submission Cap: "Repository of ALL Papers" — IgorCarron · 2026-10-06
- AI Math Breakthroughs Prompt Question: How Many Real Engineering Problems Have Math Solutions? — Afinetheorem · 2026-10-06
- AI for Math Fund adds $17.1M from XTX Markets, backing 22 projects across 30 organizations — AlexKontorovich · 2026-10-06
- 62-person blind test of 48 LLM jokes shows models are getting funnier, Astra tops with 50%+ laughs — paraschopra · 2026-10-06
- Bayesian teaching dramatically improves probabilistic reasoning in LLMs, Nature Comms paper finds — tallinzen · 2026-10-06
- LiFT loops a DiT at inference: beats DiT-XL/2 with 52% less inference compute — cgmsnoek · 2026-10-06