Loopie uses recurrent MoE layers to beat a compute-matched 30B baseline
burny_tech · x · 2026-07-21
The paper ‘Loopie’ presents a recurrent Transformer approach that is compute-efficient as well as parameter-efficient.
- The family includes a 20B-parameter MoE model with 2B active parameters and a 6B model with 0.6B active parameters.
- Instead of simply stacking more layers, each MoE layer loops twice before moving on.
- The design saves activation memory, improves throughput, and lets the model reinvest compute into capacity.
- According to the post, a compute-matched Loopie-20B-A2B beats a vanilla 30B-A3B MoE baseline and reaches gold-level IMO/IPhO performance without tools.
Related event: Loopie Looping Transformers Match Larger Models at Fraction of Cost(7 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11