Researcher argues reward is the optimization target for deeply RL-trained models

jessi_cata · x · 2026-09-13

The author takes a contrarian stance in the alignment debate over whether "reward is not the optimization target," proposing that for deeply RL-trained models, reward is the optimization target — pushing back on both the optimistic and pessimistic versions of that claim. A substantive position on the nature of RL training dynamics and what alignment research should assume about reward-seeking behavior.

Original post →

More from AGI Musings

AGI Musings channel →