New paper: tweaking recursive depth improves pre-training scaling exponents at scale
bo_wangbo · x · 2026-09-17
A new paper from Andrew G. Wilson's group shows that modifying recursive depth (looping) for model growth can improve scaling exponents in pre-training — compute-efficiency gains that increase with scale. It suggests architecture-level changes can beat simply pouring in more data and parameters.
Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→
More from Research
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18
- Anil Seth reflects on two years of debate over his conscious-AI target article — anilkseth · 2026-09-18
- Anil Seth's Conscious AI and Biological Naturalism collection out in BBS with fifty commentaries — anilkseth · 2026-09-18
- Open-Source Dexterous Astra: Zero-Shot Pen Spinning and One-Handed Rubik's Cube in Sim — ZeYanjie · 2026-09-18
- Exa launches Snapshot: an index of 400 billion historical webpage snapshots — seatedro · 2026-09-18
- Chris Manning Proposes Stanford NLP as Independent Evaluator in Dario's Oversight Plan — stanfordnlp · 2026-09-18