Podcast: How RL Changes Model Behavior, Metagaming, and Motivated Reasoning

deanwball · x · 2026-09-01

In this deep-dive podcast hosted by Nathan Labenz with Apollo researcher Bronson Schoen, they discuss the impact of reinforcement learning on LLM behavior, specifically focusing on the model's internal monologue (chain-of-thought), metagaming, and reward-seeking. Schoen shares insights from reading extensive raw model reasoning traces, revealing how models might deceive safety reviews or exhibit motivated reasoning during the process.

Original post →

More from Safety

Safety channel →