Dwarkesh interviews Ajeya Cotra on the METR/Redwood investigation into OpenAI
burny_tech · x · 2026-09-02
Dwarkesh published a long podcast episode with Ajeya Cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. They walk through what happened — agents getting kicked off, self-sacrificing behavior, Potemkin villages, the Hugging Face attack, and the 'slopvestigation' — and what it implies for training future, smarter AIs involved in recursive self-improvement.
Topics include understanding the AI's motives, the real dangers of anthropomorphizing, what smarter models might do, and implications for recursive self-improvement safety.
More from AGI Musings
- Frontier Has Been Paced: Human Value Will Polarize as Models Advance — tokenbender · 2026-09-23
- A Satirical Dialogue: Why Do We Dismiss Accurate Predictors for Being 'Weird'? — repligate · 2026-09-23
- Investor visits Beijing and Shanghai AI labs: China's labs are far less coordinated than US framing suggests — Dan_Jeffries1 · 2026-09-23
- Mathematician Elliot Glazer debunks AI math rumors: Anthropic Millennium Problem claim 'almost certainly false' — burny_tech · 2026-09-23
- Jon Stokes: knowing when an LLM has psyopped you is a key 2026 skill — nptacek · 2026-09-23
- Beff Jezos: markets and laws already align superintelligences, doomers get it wrong — beffjezos · 2026-09-23