Loopie uses recurrent MoE layers to beat a compute-matched 30B baseline
burny_tech · x · 2026-07-21
The paper ‘Loopie’ presents a recurrent Transformer approach that is compute-efficient as well as parameter-efficient.
- The family includes a 20B-parameter MoE model with 2B active parameters and a 6B model with 0.6B active parameters.
- Instead of simply stacking more layers, each MoE layer loops twice before moving on.
- The design saves activation memory, improves throughput, and lets the model reinvest compute into capacity.
- According to the post, a compute-matched Loopie-20B-A2B beats a vanilla 30B-A3B MoE baseline and reaches gold-level IMO/IPhO performance without tools.
Related event: Loopie Cyclic Transformer Matches 30B Baselines with Fractional Tokens(7 posts)→
More from Research
- Research finds memory compression makes AI agents drop safety rules and hit 59% violations — gerardsans · 2026-07-22
- DriftWorld claims a world model that runs at 30+ FPS and trains on 1–2 GPUs — du_yilun · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22