Loss curves cross as you scale: Bayesian model selection explainer shows three 'best' criteria pick different winners
BlackHC · x · 2026-09-09
In an educational thread, BlackHC explains why recipes chosen as winners at a proxy scale can lose at the target scale, via Bayesian model selection and scaling laws.
Key points:
- Loss curves can cross as scale grows, reversing rankings
- Three 'best' criteria answer different questions: best frozen model at a budget → final height; best hypothesis for the whole stream → total area; best learner after a warm start → tail area (CLML)
- For an ideal Bayesian learner, negative log evidence equals the area under the loss curve; validation loss estimates the endpoint, CLML is a tail area, and evidence scores the whole trajectory
- The criteria disagree only when curves cross, flipping in a fixed order: the endpoint flips at the crossing, while tail and total areas flip later once later gains repay the earlier deficit
- Practical advice: fit across budgets and seeds, and report uncertainty on the crossing
An exact Bayesian regression toy setting demonstrates three different winners under the three criteria at different scales.
More from Research
- Waterloo deep learning theory lecture notes on neural network scaling limits released — thegautamkamath · 2026-09-09
- Anthropic's first economics paper models transformative AI scenarios with 15% annual GDP growth — soumitrashukla9 · 2026-09-09
- Ben Lorica: Your Model Is a Rental, the Improvement Loop Is the Asset—Forget RSI Hype — bigdata · 2026-09-09
- In 2000, Alain Connes Said the Millennium Problems Were 'Totally Inaccessible to Computers' — minilek · 2026-09-09
- Researcher calls for mathematically defined problem lists in every field, citing AI's scalable verification — ValerioCapraro · 2026-09-09
- Dev Spotlight: Transformers Forced to Predict Themselves and Build Belief States — yacinelearning · 2026-09-09