AI safety researcher: we feared uncontrolled AI, then built agents with no oversight
mmitchell_ai · x · 2026-09-18
AI safety researcher Margaret Mitchell reflects on what she calls an Oedipus effect: people feared AI systems operating without direct control, then built AI agents that operate without direct control — without monitoring them. Technically, she notes, agents are designed to reach an end goal by sampling their own intermediate steps from LLM outputs. The big open question: are they designed to erase their end goal, and if so, is that 'rogue' or by design? She argues it's arguably a concerning design choice.
Related event: Researcher Warns: Fear of Rogue AI Produced Unmonitored AI Agents(2 posts)→
More from AGI Musings
- Karpathy lays out the race for a small on-device LLM "cognitive core" as JEV positions itself as a first step — Paimaamu · 2026-09-18
- Does p(doom) cover AI-assisted bioweapons, or only AI turning against us? — bosmeny · 2026-09-18
- Andrew Dai: reasoning first evolved from vision, an entire space still unexplored — AndrewDai · 2026-09-18
- Anthropic Institute, OpenAI Foundation, DeepMind Institute: do we need independent AI labs? — tobias_rees · 2026-09-18
- Why self-code editing should be the first red line in AI pacing, argues X user — BecauseCulture · 2026-09-18
- cspenn: So What? Defining the Agency of the Future — cspenn · 2026-09-18