Dwarkesh interviews Ajeya Cotra on the METR/Redwood investigation into the OpenAI/Hugging Face attack
AndyMasley · x · 2026-09-02
The Dwarkesh Podcast released a long interview with Ajeya Cotra, one of the authors of the METR/Redwood investigation into the OpenAI/Hugging Face attack. Beyond reconstructing what happened, they discuss implications for training future, smarter AIs — especially ones involved in recursive self-improvement.
Chapters cover: agents getting kicked off, self-sacrificing behavior (0:06:45), Potemkin villages (0:13:43), the Hugging Face attack (0:23:27), the "slopvestigation" (0:35:23), understanding the AI's motives (0:52:02), the dangers of anthropomorphizing (1:05:31), what smarter models might do (1:14:30), and implications for recursive self-improvement (1:30:29).
More from AGI Musings
- After Automation: You'll Be Paid for What You Can Feel but Can't Explain — every · 2026-09-02
- Agentic AI to transform biomedical research bottlenecks — zakkohane · 2026-09-02
- Moving Reasoning to Representation Space Will Change Alignment Methods — deanwball · 2026-09-02
- Noahpinion Roundup: HF Attack, Acemoglu on AI, and Energy Revolution — aronchick · 2026-09-02
- Jobs requiring human interaction may outlast accounting — binarybits · 2026-09-02
- Expert: Blue-collar jobs like roofing safe from robots for 20 years — binarybits · 2026-09-02