Scaling-law fits work down to 4M params, but point selection is unexplained

giffmana · x · 2026-09-22

The paper trains models as small as 4M params and still gets reasonable scaling-law fits. But the thread author flags a gap: the paper never discusses how it selects points for the fit — significant because the 4M model's final point is clearly off the line.

Related event: FAIR and NYU Push Scaling Laws Down to 4M Parameters, Finding Hyperparameter Tuning Is the Missing Key(7 posts)→

Original post →

More from Research

Research channel →