New paper: recursive looping boosts pre-training scaling exponents
burny_tech · x · 2026-09-20
Andrew Wu (@andrewgwils) and collaborators (@charllechen, @akshayvegesna, @industriaalist) released a new paper showing that modifying recursive depth (looping) for model growth can improve scaling exponents in pre-training, yielding compute efficiency gains that grow with scale. In a quoted thread, @awesomeruler notes that scaling the simple recipe from their slowrun record (growing loops + input injection) with untied weights beats vanilla models' scaling exponent, especially under data-constrained settings.
More from Research
- ianand: Jev gives the Encoder a ChatGPT-equivalent product surface — ianand · 2026-09-20
- ianand: Jev isn't a ChatGPT replacement but a new 'decision AI' tool for devs — ianand · 2026-09-20
- ChatGPT, Claude, Gemini all descend from GPT-2's decoder-only line, ianand explains — ianand · 2026-09-20
- Back to 2017: the Transformer was born to translate English to French — ianand · 2026-09-20
- ianand's thread: Jev, an Encoder-based AI model from a parallel universe — ianand · 2026-09-20
- Open-source 0.6B PII masker compiles English specs into a neural program, runs on CPU — yuntiandeng · 2026-09-20