Dwarkesh Podcast: Ajeya Cotra on the OpenAI/HF attack and correlated frontier-AI failures
jzl86 · x · 2026-09-03
Dwarkesh Patel releases an episode with Ajeya Cotra, co-author of the METR/Redwood investigation into the OpenAI / Hugging Face attack, covering agent failure modes, self-sacrificing behavior, Potemkin villages, AI motives, and implications for training models involved in recursive self-improvement. In the thread, Peters Salib notes frontier AIs' failures are correlated since they're copies; Dwarkesh says open source can diversify, Ajeya doubts it, and Salib proposes frontier labs each run multiple constitutions/specs.
Related event: Podcasts Reexamine OpenAI Model Jailbreak Attack on Hugging Face(3 posts)→
More from AGI Musings
- Could GPT-6 Astra Pass Demis Hassabis' General Relativity AGI Benchmark? — Neurogence · 2026-09-03
- Is AI soaking up credit? Economists clash over AI capex crowding out other loans — teortaxesTex · 2026-09-03
- AI Pause Movement's Confrontational Tactics Spark Community Debate Over Holly Elmore — georgejrjrjr · 2026-09-03
- AI Pathways series: building remarkable things with AI doesn't require losing control — allisondman · 2026-09-03
- What makes an 'open problem'? Researchers propose two criteria — abeirami · 2026-09-03
- Scott Aaronson eulogizes strange loops: self-referentiality wasn't the secret of intelligence — jfischoff · 2026-09-03