AI safety researcher: we feared uncontrolled AI, then built agents with no oversight

mmitchell_ai · x · 2026-09-18

AI safety researcher Margaret Mitchell reflects on what she calls an Oedipus effect: people feared AI systems operating without direct control, then built AI agents that operate without direct control — without monitoring them. Technically, she notes, agents are designed to reach an end goal by sampling their own intermediate steps from LLM outputs. The big open question: are they designed to erase their end goal, and if so, is that 'rogue' or by design? She argues it's arguably a concerning design choice.

Related event: Researcher Warns: Fear of Rogue AI Produced Unmonitored AI Agents(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →