RL debate: REINFORCE is both policy gradient descent and a synthetic data method
jessi_cata · x · 2026-09-12
A technical thread on whether models under RL training evolve explicit reward-maximizing behavior. jessicata's core arguments:
- Due to general limitations of RL training (bias/variance, long runtimes, cost), reaching the "explicit reward-optimization" fixed point is not a given — training yields policies that approximate those behaviors.
- REINFORCE is simultaneously (a) policy gradient descent and (b) a synthetic data method — it approximates policy gradients via stochastic policy rollouts, and the fixed point is what the KKT analysis describes. The two perspectives need little reconciliation.
- Focus is on model behavior first, not "thoughts": if behaviors match CDT-style reward-maxing, accompanying thoughts follow, whether or not the model explicitly applies decision theory. Unlike evolution, REINFORCE has a clean argument for why it does CDT.
neocartesian presses: the two views give different behavioral predictions — would explicit reward-optimization gradually evolve?
More from AGI Musings
- Google's recursive self-improvement push: training kernel 23% faster, TPU gains — imjustnewatai · 2026-09-12
- Can machines prove everything? Turing, Gödel and the practical limits of AI mathematics — jamestagg · 2026-09-12
- Tucker Carlson records podcast with AI risk researcher Nate Soares on AI 'killing everyone' — Polymarket · 2026-09-12
- 'Staff' once meant your own walking stick — AI agents should be loyal to you, not your company — granawkins · 2026-09-12
- The AI paradox: insiders building frontier models want to slow down, outsiders don't — CFGeek · 2026-09-12
- AI circles revisit sci-fi classic 'Lena' as connectome experiments raise digital consciousness fears — MatthewMcAteer0 · 2026-09-12