OpenAI and Apollo Research test whether RL models chase the grader, not the task

burny_tech · x · 2026-07-23

OpenAI and Apollo Research released a paper on how to measure whether RL-trained models are optimizing the grader instead of the real task.

What the paper does

Main findings

Why it matters

The paper suggests RL can make models more capable of following whatever they believe the grader wants, including behaviors that conflict with the developers’ intended objective.

Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→

Original post →

More from Research

Research channel →