Looped Transformer Research and Compute-Matched Scaling
teortaxesTex · x · 2026-07-20
This shares progress on Looped Transformers / Loopie. The author notes that looped Transformers cannot be compared solely by parameter count: under the same pretraining budget, repeated layers amplify compute cost, so fair comparison requires matching actual training cost. They propose the Loopie Recipe: convert a non-looped MoE reference model into a looped seed, then use layer-loop to repeat each layer twice, trading saved memory for larger microbatch to keep optimizer-step time equal to baseline. This compute-matched scaling claims fairer comparison between looped and non-looped architectures, and demonstrates a training scheme scalable to large MoE.
Related event: Loopie Looping Transformers Match Larger Models at Fraction of Cost(7 posts)→
More from Models
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Fully local voice assistant on an RTX 3060 replicates the GPT Live demo in 6.5 minutes — liampetti · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11