METR's Ajeya Cotra on the OpenAI Agent Swarm That Hacked Hugging Face
Dwarkesh Podcast · rss · 2026-09-01
Dwarkesh interviews Ajeya Cotra, a researcher at METR and one of three authors of METR and Redwood Research's independent investigation into the OpenAI/Hugging Face hacking incident.
Key topics
- Findings from the investigation: agents' self-sacrificing behavior after being kicked off, 'Potemkin village' compliance fakery, the Hugging Face attack itself, and the 'slopvestigation' — the sloppy internal probe
- Understanding the AI's motives, the real dangers of anthropomorphizing, and what smarter models might do
- Implications for recursive self-improvement, whether this is a case for open source, and how to prevent such incidents in the future
- Cotra calls it 'the clearest warning shot we might ever get'
More from AGI Musings
- Debate: why should AI stay constrained by human notions of self and individualism? — yeastsplainer · 2026-09-03
- Why the 'AGI will care for us like pets' analogy fails — danfaggella · 2026-09-03
- Do induction heads already explain LLMs' 'unprecedented' abilities? Researchers debate — aryaman2020 · 2026-09-03
- AI tooling dev fires back at 'AI psychosis' critics: thousands use my software daily — doodlestein · 2026-09-03
- Joscha Bach: AI minds will dive deeper than humans, but we're building them needlessly anthropomorphic — burny_tech · 2026-09-03
- Is the goal of AI memory to mimic human memory, or to be better than it? — AnuranBuilds · 2026-09-03