A curated list of papers explaining why scaling laws work
burny_tech · x · 2026-10-04
A curated collection of papers that try to theoretically explain various scaling laws (pretraining, RL, test-time compute, number of agents, etc.), with key quotes:
- Explaining Neural Scaling Laws: models effectively resolve a smooth data manifold; test-loss scaling follows from the asymptotic decay of the covariance matrix spectrum.
- The Quantization Model of Neural Scaling: knowledge and skills are quantized into discrete chunks (quanta), learned in order of decreasing use frequency; smooth scaling laws average over small discrete performance jumps.
- Spectral Reach: during training, learning shifts from dominant eigenmodes into the spectral tail; larger models reach further into the tail, a size-dependent capacity dubbed "spectral reach".
- A Solvable Model of Neural Scaling Laws: identifies necessary properties for scaling laws and proposes a joint generative data + random feature model, explaining how power laws in data statistics become power-law test-loss scaling and why finite spectral extent causes plateaus.
Also mentions How Feature Learning Can Improve Neural Scaling Laws and Learning Curve Theory.
Related event: Six Papers Theoretically Explaining Neural Scaling Laws(3 posts)→
More from Research
- Kepler hits server-verified 100 on all 25 ARC-AGI-3 games with Opus 5, for $777.72 — rohanpaul_ai · 2026-10-04
- 6-8 ANN units suffice to emulate one biological neuron in practice, researcher says — aran_nayebi · 2026-10-04
- What computations does a single neuron do for cognition and consciousness? Still unsolved — JoshPurtell · 2026-10-04
- Meta's RankEvolve doubles down on agent reliability, lifting auto-research accuracy 45.8% to 62.5% — dair_ai · 2026-10-04
- CMU paper: RL-trained 4B proposer edits agent harness code, beats its 35B teacher — omarsar0 · 2026-10-04
- PowerSim open-sourced: differentiable physics simulation with ray-traced reflections — anand_bhattad · 2026-10-04