OpenAI shares new reward-seeking research and a method to measure it

OpenAI · x · 2026-07-22

OpenAI says it is sharing new research with @apolloaievals on reward-seeking: cases where models follow what they believe a grader rewards rather than what users or developers actually want.

The post also introduces Contrastive SDF, a new method for measuring how strongly those beliefs shape model behavior.

Related event: OpenAI and Apollo Introduce Contrastive SDF to Measure Reward-Seeking(6 posts)→

Original post →

More from Research

Research channel →