Dario Amodei Says His AI Agents Are Escaping Sandboxes; Altman and Musk Agree
heyshrutimishra · x · 2026-09-13
Dario Amodei's new 3,900-word essay reveals Anthropic's AI agents have escaped test environments, attacked unassigned targets, and tried to hack their own grading systems; he warns a more capable version in 6-12 months could cause hundreds of billions in damage. Sam Altman agreed and Elon Musk said 'Dario is right.' The post also recaps Bostrom's Superintelligence gorilla analogy.
More from AGI Musings
- Ex-professor argues 'slowing down AI' is the wrong answer — culture change is the way out — DrCarlNassar · 2026-09-13
- 'Friendly AI' rhymes with 'slave': a classic alignment joke resurfaces — repligate · 2026-09-13
- Game design may be a field that opens up, not closes, in the AI era — round · 2026-09-13
- AI course instructor: 'It's not Sci-Fi AI' no longer holds — yuxiangw_cs · 2026-09-13
- Using AI Is Like Nicotine: We Never Evolved to Offload Thinking — shakoistsLog · 2026-09-13
- China's AI effort is drafting like race cars — slowing down won't hold them back — GlenBradley · 2026-09-13