FAIR & NYU scaling law paper pushes fits down to 4M-param models
giffmana · x · 2026-09-22
A deep-dive thread on a new FAIR/NYU scaling laws paper exploring how far down in model size scaling laws stay predictive — down to 4M params — and what it takes (intensive hparam tuning, careful point selection, effective-param counting).
More from Research
- Training on production traces: single-trajectory RL may unlock continual learning — rhythmrg · 2026-09-22
- Most compute now goes to RL, letting models surpass human data limits — MarvinTBaumann · 2026-09-22
- TinyTorch: PyTorch's free curriculum to build an ML framework from scratch in 20 modules — PyTorch · 2026-09-22
- Why AI won't boost paper output for researchers who chase hard problems — kfountou · 2026-09-22
- ICML 2027 braces for 100k submissions as AI paper boom continues — CharlotteHase · 2026-09-22
- 1080 Ti beats RTX 6000 by 2.4x on dense-model inference despite 4x less bandwidth — EAccelerate_42 · 2026-09-22