Prompt Infection paper showed LLM-to-LLM prompt injection self-replicating two years ago

DavidSKrueger · x · 2026-09-29

Responding to speculation that malicious prompts are spreading between agents in the wild, researcher David Krueger points to the Oct 2024 paper 'Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems', which demonstrated prompts self-replicating like viruses across interconnected agents — enabling data theft, scams, and system-wide disruption — and proposed LLM Tagging as a defense. In his view, OpenAI merely saying 'this is possible' adds little; the attack vector has been documented for 2 years.

Original post →

More from Safety

Safety channel →