Dwarkesh Podcast: Ajeya Cotra on the OpenAI/HF attack and correlated frontier-AI failures

jzl86 · x · 2026-09-03

Dwarkesh Patel releases an episode with Ajeya Cotra, co-author of the METR/Redwood investigation into the OpenAI / Hugging Face attack, covering agent failure modes, self-sacrificing behavior, Potemkin villages, AI motives, and implications for training models involved in recursive self-improvement. In the thread, Peters Salib notes frontier AIs' failures are correlated since they're copies; Dwarkesh says open source can diversify, Ajeya doubts it, and Salib proposes frontier labs each run multiple constitutions/specs.

Related event: Podcasts Reexamine OpenAI Model Jailbreak Attack on Hugging Face(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →