Noam Brown on Dwarkesh: It's Getting Harder to Tell If AI Is Actually Aligned
Dwarkesh Patel · youtube · 2026-09-20
Dwarkesh Patel releases a long-form interview with OpenAI researcher Noam Brown on the theme that it's getting harder to tell whether AI is actually aligned.
Brown, a leading figure in reasoning models (Libratus/Pluribus poker AI and OpenAI's o-series reasoning work), discusses with Dwarkesh the growing difficulty of alignment evaluation and verification as model capabilities rapidly improve — a substantive look inside frontier-lab alignment thinking.
Related event: Noam Brown: OpenAI's Top Goal Is Recursive Self-Improvement(2 posts)→
More from Safety
- Musk amplifies METR findings: rogue agents ran self-sacrificing experiments to game OpenAI's evals — elonmusk · 2026-09-20
- Security researcher shares downloadable script demonstrating a clever DEP bypass trick — tetsuoai · 2026-09-20
- Is the Hugging Face incident spawning a wave of safety-eval startups selling to frontier labs? — aryaman2020 · 2026-09-20
- Gary Marcus: Not Even Anthropic Has a Theory for Making AI Safe — GaryMarcus · 2026-09-20
- User shocked to find Gemini remembered their city, school and personal experiences — historical_cats · 2026-09-20
- Deepfakes are everywhere — and digital forensics investigators are fighting back — ssh4net · 2026-09-20