Full post: Bayesian Model Selection & Scaling Laws — when proxy-scale winners flip

BlackHC · x · 2026-09-09

The full article version of the thread (Andreas Kirsch, blackhc.net). TL;DR: scaling laws tell you how models improve with scale, but not what "best" means. Bayesian model-comparison criteria expose a hidden choice: select for the best final checkpoint, the best cumulative performance while learning, or the best post-warm-start performance — corresponding to the final height, total area, and tail area of a Bayesian learner's loss curve (posterior-predictive loss, negative log marginal likelihood, negative CLML). Whenever curves cross, these criteria rank models differently, which is why proxy-scale winners can lose at target scale. Similar considerations apply to pretraining and in-context learning. The criterion should follow the downstream use case, not convention.

Illustrated with two matched runs: Run A falls fast but plateaus high; Run B starts higher and finishes lower — pick B for the final checkpoint, but classical evidence (area under the curve) may prefer A. Full post includes the exact identity, flip-order results, crossing regimes, and an interactive panel.

Related event: Bayesian Model Selection Explains Why Proxy-Scale Winners Fail at Target Scale(10 posts)→

Original post →

More from Research

Research channel →