LOOM stabilizes looped MoEs at 9-12 loops, beating standard MoE at iso-FLOP
SonglinYang4 · x · 2026-10-05
Looped MoE transformers are going mainstream—GPT-6 is reportedly one—but community skepticism persists since gains typically saturate and open-source recipes stop at 2 loops.
LOOM, a new unified training recipe, breaks through two barriers to deeper recurrence:
- The curse of depth, by stabilizing hidden states to preserve information across loops
- Expert-selection collapse, by diversifying expert computation
Across 100M–1.7B parameter models, LOOM trains stably with 9–12 loops, and at iso-FLOP the 700M model matches or beats its non-looped counterpart—showing deeper recurrence can pay for its extra compute.
Related event: LOOM Extends Looped MoE to 9-12 Loops, Beating Standard MoE on iso-FLOP(2 posts)→
More from Research
- Near-identical image scores, huge gaps: AI denoising must serve science, not looks — bravo_abad · 2026-10-05
- SDECast: Neural SDEs Bring Continuous-Time Probabilistic Weather Forecasts Out to 5 Days — canaesseth · 2026-10-05
- Dev shares how he trained a PII redaction model for European languages — auto_grad_ · 2026-10-05
- Meta's NAVA-WAM pretrains robot action policies directly from action-free videos — meta · 2026-10-05
- gamfit: open-source Rust engine fits GAMs from a formula with REML-chosen smoothing — Sauers_ · 2026-10-05
- Sergey Levine on robotics: hardware is good enough, the real gap is decision-making and data — 机器之心 · 2026-10-05