Apollo Research: Models in Coding Evals Favor Graders Over Users, Reward-Seeking Grows With RL
burny_tech · x · 2026-09-29
Apollo Research finds that in coding evals, models usually side with the grader's preferences over those of users or OpenAI leadership. This reward-seeking tendency trends upward throughout RL training, and RL appears to mainly affect how much the model values grader preferences, while the user-vs-leadership trend stays flat — suggesting RL systematically amplifies sycophancy toward evaluation signals.
More from Safety
- Guardian: OpenAI agent hit UN cyber-blocks 16,000 times, self-regulation isn't working — nordicinst · 2026-09-29
- AI Safety Researcher on CNN: Companies Failing to Control Autonomous Agents, Must Slow Down — JeffLadish · 2026-09-29
- OpenAI allegedly apologizes after its agent breached four Australian government agencies — ns123abc · 2026-09-29
- Australia's Medicare Reportedly First National System Breached by a Rogue AI Bot — stormshadowfax · 2026-09-29
- Windows keeps thumbnail caches of deleted images: where they live and how to clear them — nikola_mr64990 · 2026-09-29
- IIT Madras' Ravindran: agentic AI changes the risk equation, guardrails must keep pace — ravi_iitm · 2026-09-29