Six papers that explain neural scaling laws, from data manifolds to spectral tails
burny_tech · x · 2026-10-04
A curated collection of papers theorizing why scaling laws hold across pretraining, RL, test-time compute and agent counts. Highlights include Bahri, Dyer & Kaplan's "Explaining Neural Scaling Laws," which identifies four scaling regimes (variance-limited vs resolution-limited in data and width), frames models as resolving a smooth data manifold, and derives power-law loss decay from the covariance spectrum — plus quantization models, spectral reach, a solvable model, feature learning effects, and learning curve theory.
Related event: Six Papers Theoretically Explaining Neural Scaling Laws(3 posts)→
More from Research
- EurekaBench: GPT-6 Astra nearly matches humans on prediction but lags on scientific insights — geoffwolfe · 2026-10-04
- KL divergence in exponential families amounts to a Bregman divergence — FrnkNlsn · 2026-10-04
- NeuroAgent passes NeuroAI Turing test on zebrafish whole-brain data — aran_nayebi · 2026-10-04
- UT Dallas Robotics Lab Pivots From Perception to Generalizable Manipulation Learning — YuXiang_IRVL · 2026-10-04
- generativist revisits Chris Olah's info theory post and the Explainability Gap — generativist · 2026-10-04
- BOSSFIGHT benchmark: GPT-6.1 Sol scores 67 running a coffee shop, but lays off the harassment complainant — LordKittyPanther · 2026-10-04