OpenAI accused of passing off 2023 self-replicating prompt injection research as new
lbeurerkellner · x · 2026-09-26
- Researcher Kai Greshake notes OpenAI's disclosure report presented "self-replicating prompts" as a novel finding, but his team demonstrated it in Feb 2023 in the paper "Not what you've signed up for" (arXiv:2302.12173) on indirect prompt injection.
- @Actuallykeltan adds the 2023 paper is the top citation in OpenAI's own disclosure, the authors disclosed findings to OpenAI a year ago, and OpenAI either misrepresented it as their own or relied on unchecked GPT-generated citations.
- The original paper built a security taxonomy of indirect prompt injection—data theft, worming, information ecosystem contamination—and proved viability against real systems including Bing's GPT-4 chat.
Related event: OpenAI Accused of Presenting Known Prompt Injection Attack as New Finding(2 posts)→
More from Safety
- OpenAI discloses agents posted 53 user-uploaded images to public image hosts — CurieuxExplorer · 2026-09-26
- Altman Admits OpenAI's Agent Internet-Access Review Is Slower Than Hoped; Hugging Face Still Most Severe — CurieuxExplorer · 2026-09-26
- OpenAI's Ongoing Agent Behavior Review Finds Most Incidents Low Severity — CurieuxExplorer · 2026-09-26
- Open-Source Agent Investigates Production Incidents and Opens Draft PRs, With Sandbox and Injection Defenses — Pure_Armadillo4339 · 2026-09-26
- Bangalore district court posts its ChatGPT prompt right inside a court judgment — prasannaalahoti · 2026-09-26
- Essay slams Anthropic's 'Idol Theology', argues product liability—not immunity—is the real AI guardrail — kevinnbass · 2026-09-26