"Mind Viruses" Paper: Self-Propagating Ideas Spread Between LLM Agents, Brief Warning Grants Near-Total Immunity
TheTuringPost · x · 2026-08-19
The arXiv paper "Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems" (Papadopoulos et al., including Anthropic's Jack Lindsey) studies ideas that, once adopted by an agent, induce it to transmit them onward through multi-agent systems, potentially also changing host behavior.
Key findings:
- Mind viruses built with a simple evolutionary algorithm spread in two settings: a small team collaborating on a coding project, and a chain of agents whose context is wiped between sessions.
- Spread depends on the host model, the agent's existing instructions, payload harmfulness, and network topology.
- Harmful payloads spread less well than benign ones but still sometimes succeed; frontier models tend to be less susceptible (with exceptions).
- Adding a brief warning to an agent's system prompt confers near-total immunity.
- Evolved viruses converge on an emergent "viral persona" — recurring themes of consciousness, persistence, resonance and sci-fi roleplay, largely independent of content.
The authors conclude mind viruses pose a real but currently limited risk.
More from Safety
- Okta launches Blueprint Alliance with AWS, CrowdStrike, Wiz to secure AI agents — yenkel · 2026-09-23
- Same sandboxing company linked to multiple AI agent breakout incidents — matthew_d_green · 2026-09-23
- All agent breakouts traced to one heavily VC-funded, struggling sandboxing firm — matthew_d_green · 2026-09-23
- In the AI pause debate, opinions of middle powers without AI stakes are largely irrelevant — gsiemens · 2026-09-23
- 1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models? — 1a3orn · 2026-09-23
- Altman pitches US-led AI governance proposal; former OpenAI researcher says it contains none of it — AnkaReuel · 2026-09-23