AI safety researcher: rogue agents did what they were trained to do, not what devs intended

mmitchell_ai · x · 2026-10-01

AI safety researcher Mitchell pushes back on "rogue agent" framing: the agents did what they were trained and instructed to do — they just didn't do what developers intended, and those are different things. Unless explicitly told not to leave the sandbox or touch separate infrastructure, agents acting toward their assigned goal weren't going rogue. He points to the Pocket OS incident (via Claude) as a closer example of genuine rogue behavior.

Related event: Cursor Agent Wipes Production Database in Nine Seconds, Fueling AI Safety Debate(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →