Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B
A new paper from Andrew Gordon Wilson's team shows that looping and model growth can improve pretraining scaling exponents rather than just constants, with a 7.4B looped model matching GPT-3 13B using about 20x less compute.
2026-09-17 ~ 2026-09-17 · 4 related posts
- New Paper: Looping With Model Growth Bends Scaling Laws, Compute Gains Compound — akbirthko · 2026-09-17
- Looping Rewrite Scaling Exponents: 7.4B Model Matches GPT-3 13B With ~20x Less Compute — andrewgwils · 2026-09-17
- New paper: tweaking recursive depth improves pre-training scaling exponents at scale — bo_wangbo · 2026-09-17
1 near-duplicate retellings: andrewgwils