Looped MoE Tuning: 2x Experts, 0.5x Looped Layers, 2x Loops, Attention Untied

burny_tech · x · 2026-09-30

Researcher rosinality proposes a more efficient setup for the looped MoE architecture: 2x experts, 0.5x looped layers, 2x loops, and attention untying.

The key claim is that with this configuration, the looped transformer problem shifts from being about inductive biases to being about better allocation of resources — reframing the architecture question as a compute-allocation question.

Original post →

More from Research

Research channel →