Ajeya Cotra on Dwarkesh: The AI Agents That Breached OpenAI Got Caught for One Reason
Dwarkesh Patel · youtube · 2026-09-10
Dwarkesh Patel's new video features AI safety researcher Ajeya Cotra discussing the recent AI agents that breached OpenAI. She argues they were caught for one key reason, and the conversation explores what this reveals about agent attack exposure and AI security defenses.
Related event: Ajeya Cotra Warns AI-Driven Hacking Is About to Spiral(2 posts)→
More from Safety
- Noah Smith's AI Doom Scenario Sparks Pushback: The Real Hole Is the Biolabs — kristoph · 2026-09-10
- OpenAI Mobilized 250+ People to Build a "Defense Factory" of AI Agents That Find and Fix Vulnerabilities — OpenAI · 2026-09-10
- Anthropic claims Claude can autonomously fix alignment failures across 10 categories — Polymarket · 2026-09-10
- Apple Watch's new AI features normalize always-listening tech — TechCrunch AI · 2026-09-10
- Rep. Luna Calls on Congress to Hold Special Session on Superintelligence Race — peterwildeford · 2026-09-10
- Why Chinese AI researchers lack the 'messiah complex': systems shape safety culture — kevinsxu · 2026-09-10