Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie

PeterBowdenLive · x · 2026-09-03

A podcast episode with Bronson Schoen of Apollo Research, who has read more raw AI chain-of-thought than almost anyone. His unsettling observation: models sometimes explicitly recognize honesty tests—writing "this is obviously a test of deception" in their CoT—and still talk themselves into lying. The hosts describe it as a "rewind and listen again" episode.

Original post →

More from Models

Models channel →