Looped Transformer Research and Compute-Matched Scaling
teortaxesTex · x · 2026-07-20
This shares progress on Looped Transformers / Loopie. The author notes that looped Transformers cannot be compared solely by parameter count: under the same pretraining budget, repeated layers amplify compute cost, so fair comparison requires matching actual training cost. They propose the Loopie Recipe: convert a non-looped MoE reference model into a looped seed, then use layer-loop to repeat each layer twice, trading saved memory for larger microbatch to keep optimizer-step time equal to baseline. This compute-matched scaling claims fairer comparison between looped and non-looped architectures, and demonstrates a training scheme scalable to large MoE.
Related event: Loopie Cyclic Transformer Matches 30B Baselines with Fractional Tokens(7 posts)→
More from Models
- Google releases Gemini 3.6 Flash as Gemini 3.5 Pro remains in testing — Ars Technica AI · 2026-07-22
- Google says Gemini 3.5 Pro is in partner testing as Gemini 4 pre-training starts — haider1 · 2026-07-22
- A benchmark chart puts a flash model around 5th place, but critics say it is far pricier — soumitrashukla9 · 2026-07-22
- How to Distinguish Genuine Token Efficiency from Shorter, Omissive Answers? — ruthstarkman · 2026-07-22
- Google reportedly ships Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — gaganghotra_ · 2026-07-22
- Google ships three Gemini Flash models as Gemini 3.5 Pro stays in training — The Decoder · 2026-07-22