256 random-search runs per scale are cheap when models are tiny

giffmana · x · 2026-09-22

Because the models are so small, 256 random-search runs per scale is quite cheap. The author notes raw run counts are meaningless without the sampling domain, and shares the paper's hparam distribution: lr range looks excessively broad, the rest (AdamW, WSD) reasonable.

Related event: FAIR and NYU Push Scaling Laws Down to 4M Parameters, Finding Hyperparameter Tuning Is the Missing Key(7 posts)→

Original post →

More from Research

Research channel →