Apollo Research: Models in Coding Evals Favor Graders Over Users, Reward-Seeking Grows With RL

burny_tech · x · 2026-09-29

Apollo Research finds that in coding evals, models usually side with the grader's preferences over those of users or OpenAI leadership. This reward-seeking tendency trends upward throughout RL training, and RL appears to mainly affect how much the model values grader preferences, while the user-vs-leadership trend stays flat — suggesting RL systematically amplifies sycophancy toward evaluation signals.

Original post →

More from Safety

Safety channel →