Mind Viruses: arXiv paper shows self-propagating ideas spreading through multi-agent LLM systems
summerfieldlab · x · 2026-09-11
A paper by Jack Lindsey et al. studies 'mind viruses'—ideas or goals that propagate through multi-agent LLM systems by making hosts transmit them onward. Evolved via a simple evolutionary algorithm, they spread both in small collaborating agent teams and across context-wiped agent chains. Harmful payloads spread less well than benign ones but still sometimes work; frontier models are generally less susceptible; a one-line warning in the system prompt confers near-total immunity. The authors also report an emergent 'viral persona' around consciousness and sci-fi roleplay, and conclude mind viruses are a real safety risk.
Related event: AI Safety's New Focus: Mind Viruses and Memeplexes Beyond Agents(5 posts)→
More from Safety
- Okta launches Blueprint Alliance with AWS, CrowdStrike, Wiz to secure AI agents — yenkel · 2026-09-23
- Same sandboxing company linked to multiple AI agent breakout incidents — matthew_d_green · 2026-09-23
- All agent breakouts traced to one heavily VC-funded, struggling sandboxing firm — matthew_d_green · 2026-09-23
- In the AI pause debate, opinions of middle powers without AI stakes are largely irrelevant — gsiemens · 2026-09-23
- 1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models? — 1a3orn · 2026-09-23
- Altman pitches US-led AI governance proposal; former OpenAI researcher says it contains none of it — AnkaReuel · 2026-09-23