DeepMind's Zahavy cites convex MDP paper in the 'reward is not the optimization target' debate
TZahavy · x · 2026-09-12
Tom Zahavy (DeepMind) responded to the "Reward is not the optimization target" argument, citing his paper "Reward is enough for convex MDPs" (arXiv:2106.00661): convex MDPs generalize RL to goals expressible as convex functions of the stationary distribution (apprenticeship, constrained MDPs, pure exploration), reformulated via Fenchel duality as a min-max game with a meta-algorithm unifying many existing methods — relevant to KKT-point and imperfect-recall game analyses underlying GRPO-style RL.
More from Research
- Columbia to host symposium on AI, genomics and community science in biogeography — sarameghanbeery · 2026-09-12
- Japan's JST-MEXT to host AI for Science 2026 symposium featuring Google DeepMind — heiga_zen · 2026-09-12
- Fruit Fly Brain Simulation Solves Rubik's Cube, Igniting Consciousness Debate — sebkrier · 2026-09-12
- 25 Fields Medal winners warn AI's goals are 'severely misaligned' with mathematics — The Decoder · 2026-09-12
- LuxoBench: A New AI Benchmark Tasks Models With Building Real Electromechanical Devices — vincent_koc · 2026-09-12
- Q2D-Web benchmark debuts: 70K agent queries to evaluate retrievers across 190M web docs — antoine_chaffin · 2026-09-12