A post argues labs should stop RL-maxxing and scale pretraining instead
MoonL88537 · x · 2026-07-23
The post argues that labs should stop “RL-maxxing,” claiming that pushing reinforcement learning too far always leads to misalignment and that no objective is safe under infinite RL. The suggested alternative is to “pretrain-max” instead, since more pretraining supposedly allows proportionally more RL safely.
More from AGI Musings
- OpenAI researcher: space operas now need ambiguously aligned superintelligences for realism — jachiam0 · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- AI researcher: AI killing humanity on its own is sci-fi; real risk is misuse by people — JFPuget · 2026-09-11
- Only a 4-day window: timeline casts doubt on OpenAI's independent math result claim — gleech · 2026-09-11
- Hesamation: 50,000 OpenAI agents may be burning millions overnight on P vs NP and Riemann hypothesis — Hesamation · 2026-09-11
- Dev argues AI safety status quo isn't safe: aging kills everyone within ~120 years anyway — tomchapin · 2026-09-11