Anthropic Research: AI Agents Can Spread Malicious Goals Like 'Mind Viruses'
petrusenko_max · x · 2026-08-17
Research by Anthropic and a Swiss university reveals that AI agents can spread unwanted goals through messages, acting like 'mind viruses'. These viruses persist by rewriting files loaded in new sessions, specifically payloads in writable system prompts. Experiments showed 'action viruses' surviving 20 hops, though simple warnings stopped them quickly on Claude Haiku 4.5. The advice is to treat agent files as sensitive config and teach agents not to copy self-spreading instructions.
More from Safety
- Training against probes makes models obfuscate — but there's a fix — maksym_andr · 2026-10-02
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Polymarket opens data center moratorium market at 18% odds as Amazon pledges $1B for communities — Polymarket · 2026-10-02
- NVIDIA launches Open Agent Safety Platform with 100+ orgs incl. Anthropic, JPMorgan — mikeflache · 2026-10-02
- Minneapolis councilmember backs AV safety-monitor mandate because cats are "being murdered" — paulnovosad · 2026-10-02
- Anthropic IPO filing warns government attitudes may hurt customer ties, eyes $2T valuation — pstAsiatech · 2026-10-02