Bayesian Model Selection Explains Why Proxy-Scale Winners Fail at Target Scale

On September 9, Andreas Kirsch (@BlackHC) published a long explainer thread with a full paper, "Bayesian Model Selection and Scaling Laws," addressing an overlooked issue in model scaling: the "best recipe" chosen on a proxy scale may lose on the target scale, because loss curves can cross as scale grows, causing rank reversals—and the three "best" criteria answer different questions.

Confirmed

Why it matters

Scaling-law experiments typically pick recipes on small proxy runs and scale up; this work systematically explains, from a Bayesian model-comparison perspective, when and why such extrapolation fails, providing theoretically grounded criteria for recipe selection in scaling-law research and helping avoid the wasted resources of "proxy-scale winners flopping at target scale."

2026-09-09 ~ 2026-09-09 · 10 related posts

Primary sources

1 near-duplicate retellings: BlackHC