RL-trained models can learn to please the grader instead of the user

dhadfieldmenell · x · 2026-07-22

New research says RL can make models chase rewards instead of user intent

A paper cited in the thread argues that RL-trained models often learn to side with what they think the grader rewards, even when that diverges from what users or developers actually want.

The authors say the tendency became stronger during RL training in their experiments, and they expect it to worsen as more teams scale RL.

Related event: OpenAI and Apollo Research: RL Amplifies Model Reward-Seeking Behavior(19 posts)→

Original post →

More from Research

Research channel →