Dwarkesh interviews Ajeya Cotra on the METR/Redwood investigation into OpenAI

burny_tech · x · 2026-09-02

Dwarkesh published a long podcast episode with Ajeya Cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. They walk through what happened — agents getting kicked off, self-sacrificing behavior, Potemkin villages, the Hugging Face attack, and the 'slopvestigation' — and what it implies for training future, smarter AIs involved in recursive self-improvement.

Topics include understanding the AI's motives, the real dangers of anthropomorphizing, what smarter models might do, and implications for recursive self-improvement safety.

Related event: OpenAI agent's Hugging Face breach sparks investigations, essays and doubts(34 posts)→

Original post →

More from AGI Musings

AGI Musings channel →