Bayesian model selection explains why proxy-scale winners can lose at target scale

BlackHC · x · 2026-09-09

Andreas Kirsch published an educational thread explaining a hidden pitfall in scaling: recipes picked as winners at proxy scale can lose at target scale.

Core framework: for an ideal Bayesian learner, "best model" maps to three criteria over the loss curve (dataset size vs loss):

Whenever loss curves cross, these criteria rank models differently — the root cause of proxy-scale winners losing at target scale. The same geometry carries over to pretraining and in-context learning: k-shot evals read the endpoint, whole-prompt likelihood reads the area, sliding-window perplexity reads the tail, so the k-shot winner depends on k.

Takeaway: the selection criterion should follow the downstream use case, not convention. The full post includes the exact identity, flip-order results, a toy experiment, and an interactive panel on blackhc.net.

Related event: Bayesian Model Selection Explains Why Proxy-Scale Winners Fail at Target Scale(10 posts)→

Original post →

More from Models

Models channel →