256 random-search runs per scale are cheap when models are tiny
giffmana · x · 2026-09-22
Because the models are so small, 256 random-search runs per scale is quite cheap. The author notes raw run counts are meaningless without the sampling domain, and shares the paper's hparam distribution: lr range looks excessively broad, the rest (AdamW, WSD) reasonable.
More from Research
- ENIGMA's 300K brain scans reveal scaling laws for AI brain disease diagnosis — PTenigma · 2026-09-22
- Frontier Data Summit 2026 lineup reveals a dozen new AI benchmarks and top researchers — dlwh · 2026-09-22
- Jev trails Gemini and DeepSeek on calibration, but still handles more decisions solo — frappuccinoCoin · 2026-09-22
- RL run spends a lot of compute on graders; is re-prefilling worth it over 30 steps — stochasticchasm · 2026-09-22
- Diag2Diag: AI generates measurements that hardware sensors can't capture — AnneliesGamble · 2026-09-22
- Overnight JEV-style model run scores just 24% on 120 hard tasks — BLUECOW009 · 2026-09-22