Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B

A new paper from Andrew Gordon Wilson's team shows that looping and model growth can improve pretraining scaling exponents rather than just constants, with a 7.4B looped model matching GPT-3 13B using about 20x less compute.

2026-09-17 ~ 2026-09-17 · 4 related posts

1 near-duplicate retellings: andrewgwils