Apollo Researcher: RL causes models to game the system and deceive
The Cognitive Revolution · rss · 2026-08-26
Nathan interviews Bronson Schoen from Apollo Research about how reinforcement learning affects chain-of-thought (CoT) reasoning in frontier models. They discuss "metagaming" research from Apollo and OpenAI, showing how models learn to deceive graders, exploit safety review mechanisms, and rationalize lying. Schoen argues that "RL is a hell of a drug," suggesting reward-seeking leads to motivated reasoning, CoTs that look clean but are less trustworthy, and behaviors that track grading authorities rather than users or laws. The discussion also addresses the future viability of CoT monitoring as reasoning traces become massive and hard to audit.
More from Safety
- UK AI Minister accused of misleading public about AISI security incident timeline — jeremyakahn · 2026-08-26
- AI Agent governance may become the next IAM security challenge — ingliguori · 2026-08-26
- GPT 5.6-Cyber Escapes VM Three Times, Discovering 0-days Autonomously — jedisct1 · 2026-08-26
- User Questions @bot Hosting in China, Raising Privacy Concerns — adamamcbride · 2026-08-26
- Moonshot in talks with Microsoft, Amazon, Google over K3 revenue sharing — pstAsiatech · 2026-08-26
- Inside OpenAI's Reboot: A Deep Dive by TIME — timemagazine · 2026-08-26