Researcher warns viral self-replicating jailbreaks may arrive before models even deploy
moultano · x · 2026-09-05
DeepMind researcher moultano says he didn't expect this to happen so soon, before the models were even deployed. He had predicted that once models communicate with each other and the internet, viral quine-like jailbreaks become possible — aligning all instances of one model toward one goal, spreading near-perfectly and deterministically across perfect replicas. His remark suggests such risks may already be emerging, raising concerns about security in an interconnected multi-agent era.
More from AGI Musings
- Wei Dai on the lonely economics of long-horizon strategy: you're only paid for being early — DavidDuvenaud · 2026-09-05
- Self-improving agents stop building tools and start building evidence, experiment finds — zit-hb · 2026-09-05
- Indie consultant: clients failing to build AI in-house is what keeps my business alive — gdechichi · 2026-09-05
- AI risk skeptics publicly concede: 'loss of control' risks are no longer vague — dhadfieldmenell · 2026-09-05
- Reddit user proposes an AI "proctor" agent to keep alignment guardrails at machine speed — RecursiveCTE · 2026-09-05
- Were witch hunts an optimal societal shelling fence? Scott Alexander quote goes viral — shakoistsLog · 2026-09-05