Scaling-law fits work down to 4M params, but point selection is unexplained
giffmana · x · 2026-09-22
The paper trains models as small as 4M params and still gets reasonable scaling-law fits. But the thread author flags a gap: the paper never discusses how it selects points for the fit — significant because the 4M model's final point is clearly off the line.
More from Research
- Frontier Data Summit 2026 lineup reveals a dozen new AI benchmarks and top researchers — dlwh · 2026-09-22
- RL run spends a lot of compute on graders; is re-prefilling worth it over 30 steps — stochasticchasm · 2026-09-22
- Diag2Diag: AI generates measurements that hardware sensors can't capture — AnneliesGamble · 2026-09-22
- Overnight JEV-style model run scores just 24% on 120 hard tasks — BLUECOW009 · 2026-09-22
- RL infra detail: re-prefill over PipelineRL's KV cache reuse, batch size for GPU utilization — stochasticchasm · 2026-09-22
- New psychology paper uses social identity to explain false beliefs in AI-era information environments — steverathje2 · 2026-09-22