Why winning training recipes at proxy scale can lose at target scale
gerardsans · x · 2026-09-10
A reply to BlackHC's educational thread on Bayesian model selection and scaling laws: three criteria for "best", when they disagree under scaling, and what that means for evals. The replier argues current recipe-scaling has blindspots — it ignores token isomorphism structure, treats curation and loss as generic ML problems, and that transformers learn co-occurrence geometry under repetition. Without distribution support in the frozen pretraining landscape, scaffolding and prettier losses can't fill interpolation voids; loss only scores next-token fit on the training measure.
More from Research
- NNsight 0.8 pre-release ships new engine for near-native vLLM interpretability — gsarti_ · 2026-09-10
- Minqi Jiang: AI won't replace mathematicians, it's a new kind of telescope — MinqiJiang · 2026-09-10
- Recirculation paper: leaking late-layer activations boosts frozen Gemma3's GSM8k accuracy by 21% — s_scardapane · 2026-09-10
- Paper unveils preliminary Agent Swarm RL recipe; V4.1 teams hit 30% on ProgramBench — inductionheads · 2026-09-10
- OpenAegis doubles security-agent pass rate to 58.1% via CyberFactory trajectories — jiqizhixin · 2026-09-10
- Semantic bottleneck enables non-invasive sentence decoding from MEG signals — pnpl · 2026-09-10