Researcher argues reward is the optimization target for deeply RL-trained models
jessi_cata · x · 2026-09-13
The author takes a contrarian stance in the alignment debate over whether "reward is not the optimization target," proposing that for deeply RL-trained models, reward is the optimization target — pushing back on both the optimistic and pessimistic versions of that claim. A substantive position on the nature of RL training dynamics and what alignment research should assume about reward-seeking behavior.
More from AGI Musings
- e/acc Turns 4: From Group Chat Meme to Decentralized Anti-NGO Movement — beffjezos · 2026-09-13
- HF's Aymeric Roucher mocks 'pacing the digging': corps inflating Balrog fears to lock up mithril — AymericRoucher · 2026-09-13
- Schmidhuber Maps 38 Years of Metalearning and Recursive Self-Improvement Since 1987 — SchmidhuberAI · 2026-09-13
- AI art sucks, but designers were already forced to yield: the business-vs-art line sharpens — paul_cal · 2026-09-13
- Mathematicians grapple with OpenAI's Navier-Stokes proof: human understanding now lags the proofs — geoffwolfe · 2026-09-13
- We're in the Middle of Someone Else's Timeline, Mistaking the View for the Ending — YogeshMalik · 2026-09-13