Small-scale proxies reproduce large-scale Transformer training instabilities

stochasticchasm · x · 2026-08-28

The thread discusses the arXiv paper "Small-scale proxies for large-scale Transformer training instabilities" (Wortsman et al.). Key finding: training instabilities seen in large models (attention logit growth, output-logit/log-probability divergence) are expensive to reproduce at scale, but the authors show they also appear in small models trained at high learning rates, with large-scale mitigations equally effective there. The paper systematically studies how warm-up, weight decay, and other interventions affect sensitivity of final loss to learning rate. The discussion adds that batch size warmup didn't seem to help, and results may not transfer across architectures.

Related event: Clear GDN Architecture Diagram and Small-Scale Training Instability Study Shared(2 posts)→

Original post →

More from Research

Research channel →