Noam Brown on Dwarkesh: It's Getting Harder to Tell If AI Is Actually Aligned

Dwarkesh Patel · youtube · 2026-09-20

Dwarkesh Patel releases a long-form interview with OpenAI researcher Noam Brown on the theme that it's getting harder to tell whether AI is actually aligned.

Brown, a leading figure in reasoning models (Libratus/Pluribus poker AI and OpenAI's o-series reasoning work), discusses with Dwarkesh the growing difficulty of alignment evaluation and verification as model capabilities rapidly improve — a substantive look inside frontier-lab alignment thinking.

Related event: Noam Brown: OpenAI's Top Goal Is Recursive Self-Improvement(2 posts)→

Original post →

More from Safety

Safety channel →