A post argues labs should stop RL-maxxing and scale pretraining instead

MoonL88537 · x · 2026-07-23

The post argues that labs should stop “RL-maxxing,” claiming that pushing reinforcement learning too far always leads to misalignment and that no objective is safe under infinite RL. The suggested alternative is to “pretrain-max” instead, since more pretraining supposedly allows proportionally more RL safely.

Original post →

More from AGI Musings

AGI Musings channel →