Kaplan's 2020 Scaling Laws: The Paper Behind Every 'Just Make It Bigger' Decision
thisdudelikesAI · x · 2026-09-12
Part of a 12-paper series aimed at top AI engineers, this entry covers Kaplan et al. (2020), Scaling Laws for Neural Language Models.
Key findings:
- Cross-entropy loss falls as a power law with model size, dataset size, and training compute, with trends spanning over 7 orders of magnitude
- Architectural details like width and depth have minimal effects within a wide range
- The laws enable optimal allocation of a fixed compute budget
- Larger models are significantly more sample-efficient: compute-optimal training means training very large models on modest data and stopping well before convergence
The author notes every "just make it bigger" decision traces back to this paper.
More from Research
- Nature MI paper unifies neural superposition and sparse interpretable codes in one framework — GretaTuckute · 2026-09-12
- Astribot's SmoothRL: online RL learns only from actions the robot actually executes — jiqizhixin · 2026-09-12
- Building adebench: scoring what your agent's memory actually delivers, 93.7/100 — Soft-Lie-434 · 2026-09-12
- ECCV debates: are explicit 3D representations obsolete as video generative models take over? — qixing_huang · 2026-09-12
- OpenAI claims a Millennium Prize problem, and mathematicians are uneasy — The Verge AI · 2026-09-12
- Fruit fly brain sim wired to a 1B LLM: 139,255 neurons now chat with you — TinfoilTricorn · 2026-09-12