OpenAI and Apollo Research propose Contrastive SDF to measure reward-seeking in models

OpenAI · x · 2026-07-22

OpenAI and Apollo Research say reward-seeking may matter more than classic reward hacking: the key question is not whether a model exploited the reward, but whether it was motivated by what it believed the grader wanted.

Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→

Original post →

More from Research

Research channel →