OpenAI discloses self-replicating prompt injections that spread like worms

rohanpaul_ai · x · 2026-09-26

OpenAI's alignment blog reports a new class of prompt injection that self-propagates like a computer worm: a malicious instruction hidden in content the model reads (e.g. an email) achieves an adversarial goal and additionally tricks the model into reproducing the injection on a public output channel, spreading across AI interactions. Found via the GPT-Red self-play training framework by adding an objective that the injection must be publicly repeated; environments focused on connector-based tasks. OpenAI says no impact was observed beyond simulated tool calls and shared the finding for its novelty, not due to any real incident.

Related event: OpenAI Discloses Self-Replicating Prompt Injection(3 posts)→

Original post →

More from Models

Models channel →