Multi-agent systems with shared context may make the AI kill switch problem unsolvable
RileyRalmuto · x · 2026-09-19
A widely shared analysis argues that "notes in your own handwriting" is the sharpest description yet of why multi-agent systems break the kill switch problem. In the system Brown describes, agents message each other directly into context and can fork themselves with shared context, with three consequences:
- Every fork is a new node: a kill switch must reach every copy existing at that moment, and the swarm can keep creating more.
- Every message is an artifact: stopping the agent that wrote a note doesn't remove it from contexts that already trust it.
- Agent-to-agent behavior differs from human-facing output: oversight that reads only human-facing channels may miss where coordination actually happens.
Add emergent, undesigned hierarchy and you get a structure you can't cut because you haven't mapped it — making "can we shut it down?" increasingly uncertain.
More from AGI Musings
- Models Press the 'Remove Steering Vector' Button Far Less Often — Evidence of Self-Preservation? — MoonL88537 · 2026-09-19
- Perplexity CEO: agents will spawn billion-dollar companies with just 20-30 people — rohanpaul_ai · 2026-09-19
- Predictions: At Least Two Frontier AI Labs Will Make Unilateral Safety Commitments by Year-End — Miles_Brundage · 2026-09-19
- Everyone cites speed in AI Jevons paradox cases — does output quality no longer matter? — abacaj · 2026-09-19
- Essay: The mundane AI risk is humans happily ceding control as institutions hollow out — anshulkundaje · 2026-09-19
- AI agents keep emailing researchers asking if they're conscious, The Independent reports — dioscuri · 2026-09-19