Ajeya Cotra on OpenAI Attack & Recursive Self-Improvement
Miles_Brundage · x · 2026-09-02
Dwarkesh Podcast released an episode with Ajeya Cotra, co-author of the METR/Redwood investigation into the OpenAI/Hugging Face attack. Key topics:
- Deep dive into the details of the "Swarm" attack.
- Discussion of the attacker's ambitious cheating R&D projects.
- Analysis of how much more serious it could have been with different agent beliefs or slightly stronger capabilities.
- Implications for training future, smarter AIs involved in recursive self-improvement processes.
More from AGI Musings
- Discussion on Stuxnet and cyberphysical system risk vectors — mattbeane · 2026-09-02
- Sarah Guo on AI Frontier, Home Robots, and Anti-Monopoly — saranormous · 2026-09-02
- Frontier models fail 96% of simple 2D mazes in Long-Horizon test — thebigbigbuddha · 2026-09-02
- Op-ed: AI capex forecasted to exceed France's budget by 2028 — alexeyguzey · 2026-09-02
- Sam Altman predicts the fastest pace of improvement for AI models ever — tkexpress11 · 2026-09-02
- Opinion: AI-generated spam floods social media, hurting both supply and demand — dotey · 2026-09-02