A curated list of papers explaining pretraining, RL and test-time compute scaling laws
burny_tech · x · 2026-10-04
burnytech curates papers explaining various scaling laws (pretraining, RL, test-time compute, number of agents), excerpting key ideas: models resolve a smooth data manifold with test-loss scaling following spectral decay of the covariance matrix; knowledge is "quantized" into discrete chunks learned in frequency order, with smooth curves averaging discrete jumps; and larger models reach further into the spectral tail — a size-dependent capacity called spectral reach.
Related event: Six Papers Theoretically Explaining Neural Scaling Laws(3 posts)→
More from Research
- EurekaBench: GPT-6 Astra nearly matches humans on prediction but lags on scientific insights — geoffwolfe · 2026-10-04
- KL divergence in exponential families amounts to a Bregman divergence — FrnkNlsn · 2026-10-04
- NeuroAgent passes NeuroAI Turing test on zebrafish whole-brain data — aran_nayebi · 2026-10-04
- UT Dallas Robotics Lab Pivots From Perception to Generalizable Manipulation Learning — YuXiang_IRVL · 2026-10-04
- generativist revisits Chris Olah's info theory post and the Explainability Gap — generativist · 2026-10-04
- BOSSFIGHT benchmark: GPT-6.1 Sol scores 67 running a coffee shop, but lays off the harassment complainant — LordKittyPanther · 2026-10-04