Counterintuitive finding: bigger models make good hparams easier to find

giffmana · x · 2026-09-22

From the random-search result distributions, the paper concludes (surprisingly) that good hyperparameters are easier to find at larger scale, and very hard at the smallest scale. The author argues this needs qualification: it doesn't mean training gets simpler with scale — instabilities grow beyond the paper's range; what it does show is that below 64M you need pretty intensive tuning.

Related event: FAIR and NYU Push Scaling Laws Down to 4M Parameters, Finding Hyperparameter Tuning Is the Missing Key(7 posts)→

Original post →

More from Research

Research channel →