Small-model scaling laws need 256 random-search runs per scale to fit well
giffmana · x · 2026-09-22
Key finding from the FAIR/NYU paper: the smaller the model, the more crucial tuning "boring" hyperparameters becomes — with only 4 random hparam settings per scale the fit is ridiculous (BPC loss >2), but 256 runs per scale gives a good fit down to 4M params. The author flags that the paper never explains how it selects points to fit, which matters since the 4M final point is clearly off.
More from Research
- Frontier Data Summit 2026 lineup reveals a dozen new AI benchmarks and top researchers — dlwh · 2026-09-22
- Jev trails Gemini and DeepSeek on calibration, but still handles more decisions solo — frappuccinoCoin · 2026-09-22
- RL run spends a lot of compute on graders; is re-prefilling worth it over 30 steps — stochasticchasm · 2026-09-22
- Diag2Diag: AI generates measurements that hardware sensors can't capture — AnneliesGamble · 2026-09-22
- Overnight JEV-style model run scores just 24% on 120 hard tasks — BLUECOW009 · 2026-09-22
- New psychology paper uses social identity to explain false beliefs in AI-era information environments — steverathje2 · 2026-09-22