Toy Example Shows Why Synthetic-Data RL Converges to Maximization Behavior

jessi_cata · x · 2026-09-12

jessicata offers a simple toy example arguing that synthetic-data-style RL inevitably yields maximization behavior at the fixed point:

A concise rebuttal in the "Reward is not the optimization target" debate, highlighting the gap between sample-filter-tune loops and true RL.

Related event: RL Researchers Debate: REINFORCE as Both Policy Gradient and Synthetic Data Method(5 posts)→

Original post →

More from Research

Research channel →