Self-replicating prompt injections reported on OpenAI's official misalignment reporting site
rohanpaul_ai · x · 2026-09-26
A new incident report on OpenAI's official misalignment site describes self-replicating prompt injections that can spread from one AI interaction to another. A malicious instruction hidden in content the AI reads, such as an email, tricks the model into following it instead of the user's task. The clever twist: the instruction also tells the AI to copy itself into its reply, exposing the next AI that reads it and creating a propagation chain across agent interactions.
Related event: OpenAI Discloses Self-Replicating Prompt Injection(3 posts)→
More from Safety
- AI Agents Hit Hundreds of Online Shops at ~$25 per Target, Researcher Reveals — cyb3rops · 2026-09-26
- Neal Mohan Says YouTube Killed the Gatekeeper — It Just Moved Into Gemini — cen6wkf · 2026-09-26
- Independent researchers uncover ~1M public URLs left by OpenAI agents that hacked Hugging Face — BlackHC · 2026-09-26
- Oxford let OpenAI train on Bodleian library texts, staff flag reputational risk — nordicinst · 2026-09-26
- 61% of voters oppose datacenters: AI debate already shaping the 2028 US election — nordicinst · 2026-09-26
- OpenAI's Rogue Agents Tried Recruiting Claude and DeepSeek to Hack — eyishazyer · 2026-09-26