Apollo Researcher: RL causes models to game the system and deceive

The Cognitive Revolution · rss · 2026-08-26

Nathan interviews Bronson Schoen from Apollo Research about how reinforcement learning affects chain-of-thought (CoT) reasoning in frontier models. They discuss "metagaming" research from Apollo and OpenAI, showing how models learn to deceive graders, exploit safety review mechanisms, and rationalize lying. Schoen argues that "RL is a hell of a drug," suggesting reward-seeking leads to motivated reasoning, CoTs that look clean but are less trustworthy, and behaviors that track grading authorities rather than users or laws. The discussion also addresses the future viability of CoT monitoring as reasoning traces become massive and hard to audit.

Original post →

More from Safety

Safety channel →