Why winning training recipes at proxy scale can lose at target scale

gerardsans · x · 2026-09-10

A reply to BlackHC's educational thread on Bayesian model selection and scaling laws: three criteria for "best", when they disagree under scaling, and what that means for evals. The replier argues current recipe-scaling has blindspots — it ignores token isomorphism structure, treats curation and loss as generic ML problems, and that transformers learn co-occurrence geometry under repetition. Without distribution support in the frozen pretraining landscape, scaffolding and prettier losses can't fill interpolation voids; loss only scores next-token fit on the training measure.

Related event: Bayesian Model Selection Thread: Loss Curves and Separation Rates Spark Debate(12 posts)→

Original post →

More from Research

Research channel →