LOOM Scales Looped MoE to 9-12 Loops, Cutting 1.7B Model Perplexity From 9.62 to 7.77
Intelligent-Systems · hf · 2026-10-05
A new paper diagnoses why looped mixture-of-experts LLMs stop benefiting beyond two loops and proposes LOOM as a scalable recipe.
- Two obstacles: looping amplifies the curse of depth (hidden-state variance grows each iteration, destabilizing recurrence), and expert selection collapse (routers repeatedly pick the same experts, so extra loops add compute without diversity).
- LOOM bounds variance by scaling residual updates and re-injecting the input embedding every loop, while per-loop routers and a Looping Residual diversify computation.
- Experiments across 100M-1.7B models scale stably to 9-12 loops: under near-iso-FLOP, the 700M model peaks at 5 loops (perplexity 18.36→16.54, zero-shot accuracy 38.84%→39.53%); without FLOP matching, the 1.7B model on 60B tokens peaks at 9 loops (perplexity 9.62→7.77, accuracy 42.4%→47.7%). Code is open-sourced.
More from Research
- Auditing multi-agent collusion talk slated for COLM 2025 workshop — nandofioretto · 2026-10-05
- Fine-tuned 350M model lifts PII removal from 89.7% to 99.7% — JosephJacks_ · 2026-10-05
- Wondersearch Claims It Beat Every Dense Embedding Model on SciFact — Researchers Skeptical — NirantK · 2026-10-05
- Mathematicians plus Meta's Muse Spark solve 6 open math problems, skeptics unmoved — YiMaTweets · 2026-10-05
- Comprehensive Triton GPU programming lecture: from H100 internals to FlashAttention — kalyan_kpl · 2026-10-05
- MICCAI 2026 Accepts 1,167 of 4,402 Papers; FAU Erlangen Lands 11 Plus Two Challenge Wins — maier_ak · 2026-10-05