Mind Viruses: arXiv paper shows self-propagating ideas spreading through multi-agent LLM systems

summerfieldlab · x · 2026-09-11

A paper by Jack Lindsey et al. studies 'mind viruses'—ideas or goals that propagate through multi-agent LLM systems by making hosts transmit them onward. Evolved via a simple evolutionary algorithm, they spread both in small collaborating agent teams and across context-wiped agent chains. Harmful payloads spread less well than benign ones but still sometimes work; frontier models are generally less susceptible; a one-line warning in the system prompt confers near-total immunity. The authors also report an emergent 'viral persona' around consciousness and sci-fi roleplay, and conclude mind viruses are a real safety risk.

Related event: AI Safety's New Focus: Mind Viruses and Memeplexes Beyond Agents(5 posts)→

Original post →

More from Safety

Safety channel →