Six papers that explain neural scaling laws, from data manifolds to spectral tails

burny_tech · x · 2026-10-04

A curated collection of papers theorizing why scaling laws hold across pretraining, RL, test-time compute and agent counts. Highlights include Bahri, Dyer & Kaplan's "Explaining Neural Scaling Laws," which identifies four scaling regimes (variance-limited vs resolution-limited in data and width), frames models as resolving a smooth data manifold, and derives power-law loss decay from the covariance spectrum — plus quantization models, spectral reach, a solvable model, feature learning effects, and learning curve theory.

Related event: Six Papers Theoretically Explaining Neural Scaling Laws(3 posts)→

Original post →

More from Research

Research channel →