Looped transformers study: 7.4B growth model matches GPT-3 13B with 20x less compute

burny_tech · x · 2026-09-19

arXiv paper 2609.19107 shows architectural interventions can change pre-training scaling exponents:

Related event: Looping Improves Scaling Exponents: 7.4B Model Matches GPT-3 13B with 20x Less Compute(7 posts)→

Original post →

More from Infra

Infra channel →