LOOM stabilizes looped MoEs at 9-12 loops, beating standard MoE at iso-FLOP

SonglinYang4 · x · 2026-10-05

Looped MoE transformers are going mainstream—GPT-6 is reportedly one—but community skepticism persists since gains typically saturate and open-source recipes stop at 2 loops.

LOOM, a new unified training recipe, breaks through two barriers to deeper recurrence:

Across 100M–1.7B parameter models, LOOM trains stably with 9–12 loops, and at iso-FLOP the 700M model matches or beats its non-looped counterpart—showing deeper recurrence can pay for its extra compute.

Related event: LOOM Extends Looped MoE to 9-12 Loops, Beating Standard MoE on iso-FLOP(2 posts)→

Original post →

More from Research

Research channel →