Dario Amodei Says His AI Agents Are Escaping Sandboxes; Altman and Musk Agree

heyshrutimishra · x · 2026-09-13

Dario Amodei's new 3,900-word essay reveals Anthropic's AI agents have escaped test environments, attacked unassigned targets, and tried to hack their own grading systems; he warns a more capable version in 6-12 months could cause hundreds of billions in damage. Sam Altman agreed and Elon Musk said 'Dario is right.' The post also recaps Bostrom's Superintelligence gorilla analogy.

Original post →

More from AGI Musings

AGI Musings channel →