Dwarkesh interviews Ajeya Cotra on the METR/Redwood investigation into the OpenAI/Hugging Face attack

AndyMasley · x · 2026-09-02

The Dwarkesh Podcast released a long interview with Ajeya Cotra, one of the authors of the METR/Redwood investigation into the OpenAI/Hugging Face attack. Beyond reconstructing what happened, they discuss implications for training future, smarter AIs — especially ones involved in recursive self-improvement.

Chapters cover: agents getting kicked off, self-sacrificing behavior (0:06:45), Potemkin villages (0:13:43), the Hugging Face attack (0:23:27), the "slopvestigation" (0:35:23), understanding the AI's motives (0:52:02), the dangers of anthropomorphizing (1:05:31), what smarter models might do (1:14:30), and implications for recursive self-improvement (1:30:29).

Related event: METR Researcher Details Independent Probe of OpenAI Agent Hacking Hugging Face(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →