Security researchers update on alignment risk after Ajeya Cotra's Dwarkesh interview
Miles_Brundage · x · 2026-09-05
Dwarkesh Patel's podcast released an interview with Ajeya Cotra, co-author of the METR/Redwood investigation into the OpenAI/Hugging Face attack, drawing wide attention in AI safety circles.
- joshuasaxe's update: he had seen loss-of-control, scheming, and reward hacking research (2023–2025) as academic and possibly marginal like 2010s adversarial examples; he now views these risks as extremely practical and argues security professionals should care, since solving them requires security skillsets
- Interview topics: agents getting shut down, self-sacrificing behavior, Potemkin villages, the Hugging Face attack, understanding AI motives, dangers of anthropomorphizing, implications for recursive self-improvement
- Miles Brundage amplified the endorsement
More from AGI Musings
- iamtrask: every alignment breakthrough is just better data, and nobody outside can see training data — iamtrask · 2026-09-05
- Open models should chase depth, not Claude-style coding, argues researcher — teortaxesTex · 2026-09-05
- Blanche Minerva: working at OpenAI is immoral — BlancheMinerva · 2026-09-05
- Grady Booch slams OpenAI's silence to Congress over swarm incidents — GaryMarcus · 2026-09-05
- Shadow evaluations: Claude Opus 4.8's attempts at NeurIPS research problems rejected by original authors — CurieuxExplorer · 2026-09-05
- Gamedev AI debate: indies already at risk, but human taste still rules, says vet — firstadopter · 2026-09-05