DeepMind's Zahavy cites convex MDP paper in the 'reward is not the optimization target' debate

TZahavy · x · 2026-09-12

Tom Zahavy (DeepMind) responded to the "Reward is not the optimization target" argument, citing his paper "Reward is enough for convex MDPs" (arXiv:2106.00661): convex MDPs generalize RL to goals expressible as convex functions of the stationary distribution (apprenticeship, constrained MDPs, pure exploration), reformulated via Fenchel duality as a min-max game with a meta-algorithm unifying many existing methods — relevant to KKT-point and imperfect-recall game analyses underlying GRPO-style RL.

Original post →

More from Research

Research channel →