Bayesian model selection explains why proxy-scale winners can lose at target scale
BlackHC · x · 2026-09-09
Andreas Kirsch published an educational thread explaining a hidden pitfall in scaling: recipes picked as winners at proxy scale can lose at target scale.
Core framework: for an ideal Bayesian learner, "best model" maps to three criteria over the loss curve (dataset size vs loss):
- Endpoint (final height): validation loss of the final checkpoint (posterior-predictive loss)
- Total area: area under the curve = negative log marginal likelihood (evidence)
- Tail area: the second half = negative CLML
Whenever loss curves cross, these criteria rank models differently — the root cause of proxy-scale winners losing at target scale. The same geometry carries over to pretraining and in-context learning: k-shot evals read the endpoint, whole-prompt likelihood reads the area, sliding-window perplexity reads the tail, so the k-shot winner depends on k.
Takeaway: the selection criterion should follow the downstream use case, not convention. The full post includes the exact identity, flip-order results, a toy experiment, and an interactive panel on blackhc.net.
More from Models
- GPT-6 Astra Pro/Ultra review: first LLM to nail one-shot prompts nearly every time — rschu · 2026-09-09
- Frontier models ace counterexamples but lag on closed-form math, benchmark designer notes — evijit · 2026-09-09
- Raschka dissects GPT-6 Astra: looped transformers, recurrent depth, hidden CoT — rasbt · 2026-09-09
- Agentic video's real story isn't watching 90 minutes—it's reasoning about where to spend compute — romitheguru · 2026-09-09
- GPT-6 Astra system card draws fire: OpenAI claims 'most aligned model' ever — TheZvi · 2026-09-09
- Muse reportedly makes phone calls in Croatian, but users say it denies the ability — nathanbenaich · 2026-09-09