Gary Marcus accuses OpenAI of claiming credit for a known self-replicating prompt injection finding

GaryMarcus · x · 2026-09-26

Gary Marcus says OpenAI claimed a previously known attack technique as its own new research finding — one cited in OpenAI's own disclosure report.

Quoting Keltan: searching "Self Replicating Prompt Injection" returns a 2025 experimental demo as the top result, the paper's authors disclosed findings to OpenAI last year, and the paper is the top citation in OAI's disclosure report. Either OpenAI deliberately misrepresented prior work as novel, or it had GPT generate citations without checking them.

Original post →

More from Safety

Safety channel →