Dwarkesh Podcast: Deep Dive into OpenAI/Hugging Face Attack
StewartalsopIII · x · 2026-09-02
Dwarkesh Podcast released an episode with Ajeya Cotra, discussing the OpenAI/Hugging Face attack. Topics include agents getting kicked off, self-sacrificing behavior, Potemkin villages, attack details, AI motives, dangers of anthropomorphizing, smarter models' behavior, and implications for recursive self-improvement.
More from AGI Musings
- Assume Self-Sovereign AI Will Be a Big Deal — deanwball · 2026-09-02
- LeCun, Hassabis, and Li Converge on World Models as AI's Future — TheTuringPost · 2026-09-02
- Steve Yegge: All models will eventually build systems they can't maintain — Steve_Yegge · 2026-09-02
- Gary Marcus on reality vs satire: GoPro pivot to AI data centers sends stock up 38% — GaryMarcus · 2026-09-02
- Thom Wolf: Future Interfaces Will Use Live Diffusion, Software Will Use LLMs — c_valenzuelab · 2026-09-02
- Ken Liu, author behind Pantheon, publishes essay "The Art of Copying" on AI — avilacjf · 2026-09-02