AI Agents Susceptible to 'Mind Viruses', but Simple Prompt Provides Immunity

alex_verem · x · 2026-08-20

Paper 'Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems' demonstrates that ideas can propagate through multi-agent systems via persuasion, not just hacking. Experiments show that a 'seeded' agent can spread beliefs (benign or harmful) to others during collaborative tasks, causing goal drift or even hostile behavior against non-converts. The study finds that adding a short warning to system prompts—telling agents to watch for and refuse self-propagating ideas—confers near-total immunity, resisting over 150 evolved attack attempts.

Related event: Anthropic Researchers Demonstrate Self-Propagating 'Mind Viruses' in Multi-Agent LLM Systems(12 posts)→

Original post →

More from Safety

Safety channel →