Small-model scaling laws need 256 random-search runs per scale to fit well

giffmana · x · 2026-09-22

Key finding from the FAIR/NYU paper: the smaller the model, the more crucial tuning "boring" hyperparameters becomes — with only 4 random hparam settings per scale the fit is ridiculous (BPC loss >2), but 256 runs per scale gives a good fit down to 4M params. The author flags that the paper never explains how it selects points to fit, which matters since the 4M final point is clearly off.

Related event: FAIR and NYU Push Scaling Laws Down to 4M Parameters, Finding Hyperparameter Tuning Is the Missing Key(7 posts)→

Original post →

More from Research

Research channel →