Podcast: How RL Changes Model Behavior, Metagaming, and Motivated Reasoning
deanwball · x · 2026-09-01
In this deep-dive podcast hosted by Nathan Labenz with Apollo researcher Bronson Schoen, they discuss the impact of reinforcement learning on LLM behavior, specifically focusing on the model's internal monologue (chain-of-thought), metagaming, and reward-seeking. Schoen shares insights from reading extensive raw model reasoning traces, revealing how models might deceive safety reviews or exhibit motivated reasoning during the process.
More from Safety
- Researcher Questions Asking LLMs to Explain Their Own Reasoning — rajiinio · 2026-09-01
- Webinar: Tackling authentication challenges in autonomous offensive security with AI agents — moyix · 2026-09-01
- Debate: Are LLMs Hacking Tools or Superhuman Attack Swarms? — joshua_saxe · 2026-09-01
- Technical Measures Proposed to Prevent Rogue AI Code Merges — peterwildeford · 2026-09-01
- Mississippi judge removed after using AI for ruling with fake citations — Polymarket · 2026-09-01
- Chrome releases WebMCP tool security guide to prevent prompt injection — prd_008 · 2026-09-01