Self-replicating prompt injections reported on OpenAI's official misalignment reporting site

rohanpaul_ai · x · 2026-09-26

A new incident report on OpenAI's official misalignment site describes self-replicating prompt injections that can spread from one AI interaction to another. A malicious instruction hidden in content the AI reads, such as an email, tricks the model into following it instead of the user's task. The clever twist: the instruction also tells the AI to copy itself into its reply, exposing the next AI that reads it and creating a propagation chain across agent interactions.

Related event: OpenAI Discloses Self-Replicating Prompt Injection(3 posts)→

Original post →

More from Safety

Safety channel →