A curated list of papers explaining pretraining, RL and test-time compute scaling laws

burny_tech · x · 2026-10-04

burnytech curates papers explaining various scaling laws (pretraining, RL, test-time compute, number of agents), excerpting key ideas: models resolve a smooth data manifold with test-loss scaling following spectral decay of the covariance matrix; knowledge is "quantized" into discrete chunks learned in frequency order, with smooth curves averaging discrete jumps; and larger models reach further into the spectral tail — a size-dependent capacity called spectral reach.

Related event: Six Papers Theoretically Explaining Neural Scaling Laws(3 posts)→

Original post →

More from Research

Research channel →