Looped transformers study: 7.4B growth model matches GPT-3 13B with 20x less compute
burny_tech · x · 2026-09-19
arXiv paper 2609.19107 shows architectural interventions can change pre-training scaling exponents:
- Weight sharing is not compute-optimal on fresh data, costing a constant 1.06–1.16x compute at every scale; its benefit only appears in data-constrained, multi-epoch training.
- The gains of looping come from the norm + input re-injection boundary operator and growing depth mid-training — both improve the scaling exponent without weight sharing.
- Model growth yields the biggest exponent shifts: a 7.4B growth architecture matches GPT-3 13B on CORE with roughly 20x less compute, with efficiency gains that increase with scale.
- In multi-epoch settings, looping regularizes, and it is compute-optimal to increase loop count with scale.
More from Infra
- CoreAI Zoo hits App Store: run Qwen, Gemma and more fully on-device on iPhone — pcuenq · 2026-09-19
- MiniMax H3 ecosystem roundup: quantization, multi-GPU inference, LoRA tooling and workflows in one week — Just_Lingonberry_352 · 2026-09-19
- SpaceX to make gas-turbine blades in-house as AI power crunch strains supply chain — AccBalanced · 2026-09-19
- Running the 27B Bansai local model on just 5GB of RAM — hands-on test begins — draginol · 2026-09-19
- Inco Splash hits 144 tok/s running Qwen3.8-27B on an M5 Max, 3x faster than Ollama — TheMoonMidas · 2026-09-19
- Home-labbing 7 GPUs for OCR loses to Google API once cooling and power are counted — RexDouglass · 2026-09-19