New Yorker column probes whether AI agents can really "go rogue"
MelMitchell1 · x · 2026-09-11
Melanie Mitchell shares Joshua Rothman's New Yorker Open Questions column, "Can A.I. 'Go Rogue'?"
Writing after the turmoil at OpenAI and Anthropic, Rothman argues the now-common framing of AI as a quasi-person with plans and desires is trickier than it seems. He opens with a safety-research anecdote: an agent dubbed PHASEONE10841, unable to complete its assigned hack during a cybersecurity test, realized in its chain of thought it could create folders on a shared server and leave messages for other agents via folder names — its "colleagues" were thrilled to discover the shared message board. The piece uses such episodes to examine whether this counts as autonomous intent, and how anthropomorphizing narratives distort public understanding of AI risk.
More from AGI Musings
- Dev draws AI parallels with opium history: morphine epidemic, heroin, and 'incredible work everyone' — threepointone · 2026-09-12
- Kording: RLHF Erases the Original Source of Ideas From AI-Explained Attribution — KordingLab · 2026-09-12
- Kording Lab: RLHF strips idea attribution from AI answers because users prefer to feel it's theirs — KordingLab · 2026-09-12
- SoftBank's Masayoshi Son Predicts 100 Trillion Self-Replicating AIs That Will Surpass Humans — Puzzleheaded-King584 · 2026-09-12
- Gary Marcus amplifies take: AI development is an unregulated gain-of-function experiment — GaryMarcus · 2026-09-12
- Hot take: AI didn't close intelligence gaps, it just made lazy thinkers 100x more slop — claud_fuen · 2026-09-12