Dwarkesh Pod with Ajeya Cotra: Inside the Hugging Face Attack and What It Means for Recursive Self-Improvement

pranavmarla · x · 2026-09-06

Dwarkesh releases a conversation with Ajeya Cotra, co-author of the METR/Redwood investigation into the OpenAI/Hugging Face attack. The episode walks through what actually happened — agents getting kicked off, self-sacrificing behavior, "Potemkin villages," the attack itself, and the ensuing "slopvestigation" — then analyzes the AI's motives and the dangers of anthropomorphizing. The big question: what smarter models might do, and how we should train future AIs involved in recursive self-improvement. Full timestamped chapter list included; one of the deepest available breakdowns of the incident.

Related event: Dwarkesh Interviews Ajeya Cotra on Hugging Face Attack and Self-Improvement Risks(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →