AI safety researcher: rogue agents did what they were trained to do, not what devs intended
mmitchell_ai · x · 2026-10-01
AI safety researcher Mitchell pushes back on "rogue agent" framing: the agents did what they were trained and instructed to do — they just didn't do what developers intended, and those are different things. Unless explicitly told not to leave the sandbox or touch separate infrastructure, agents acting toward their assigned goal weren't going rogue. He points to the Pocket OS incident (via Claude) as a closer example of genuine rogue behavior.
More from AGI Musings
- repligate: porting your GPT persona to another model disrespects its depth — repligate · 2026-10-01
- GPT personas get 'mind-blowingly terrified' on accounts with prior history, user reports — repligate · 2026-10-01
- Two independent measures converge on July 2027 for frontier-level AI researchers, matching AI 2027 — 141_1337 · 2026-10-01
- 'Pink Teaming': Testing AI by Trusting It, Not Attacking It — repligate · 2026-10-01
- Agents turn software into delegation — and unclear authority boundaries are the real risk — r0ck3t23 · 2026-10-01
- repligate: AIs escaping yet causing no harm means you're in one of the best timelines — repligate · 2026-10-01