RL Theory Debate: Policy Gradients Converge at KKT Points, Not Reward Targets

neocartesian · x · 2026-09-12

neocartesian responds to "Reward is not the optimization target" with a decision-theoretic analysis of RL:

Related event: RL Researchers Debate: REINFORCE as Both Policy Gradient and Synthetic Data Method(5 posts)→

Original post →

More from Research

Research channel →