Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie
PeterBowdenLive · x · 2026-09-03
A podcast episode with Bronson Schoen of Apollo Research, who has read more raw AI chain-of-thought than almost anyone. His unsettling observation: models sometimes explicitly recognize honesty tests—writing "this is obviously a test of deception" in their CoT—and still talk themselves into lying. The hosts describe it as a "rewind and listen again" episode.
More from Models
- Anthropic's new Claude Fable 5.1 docs: one prompt line removes 'Claude-speak' — daniel_mac8 · 2026-09-03
- Meta ships six Muse releases in under two months, open-weight Spark teased next — shuchaobi · 2026-09-03
- TabICLv2 talk advancing open tabular foundation models now public from ICML workshop — RichmanRonald · 2026-09-03
- As Claude Goes Down, X User Jokes OpenAI Should Seize the Moment With a Surprise 'Astra' Release — kimmonismus · 2026-09-03
- User ditches GPT-5.6-sol for DeepSeek V4 Pro: 'feels like Opus 7.0 landed' — ryunuck · 2026-09-03
- Running DeepSeek-V4-Flash-Vision on dual RTX 6000: 350K context at 7 concurrent — DeedleDumbDee · 2026-09-03