Gary Marcus accuses OpenAI of claiming credit for a known self-replicating prompt injection finding
GaryMarcus · x · 2026-09-26
Gary Marcus says OpenAI claimed a previously known attack technique as its own new research finding — one cited in OpenAI's own disclosure report.
Quoting Keltan: searching "Self Replicating Prompt Injection" returns a 2025 experimental demo as the top result, the paper's authors disclosed findings to OpenAI last year, and the paper is the top citation in OAI's disclosure report. Either OpenAI deliberately misrepresented prior work as novel, or it had GPT generate citations without checking them.
More from Safety
- Oxford quietly let OpenAI train on the Bodleian Library's books — ross2000 · 2026-09-26
- skill-audit: Open-source pre-install auditor targets rising agent skill supply-chain risk — masiha97 · 2026-09-26
- Farmer's photo exposes 62 unpermitted gas generators powering Microsoft AI data center, $1.1M fine — mkheck · 2026-09-26
- Malware now asks commercial AI models to vote on its next attack move — TechNadu · 2026-09-26
- Fact-Checking the NeurIPS-Pangram Story: v3.3.2 Flagged 14.9% of Human-Written Reviews as AI — tughanbulut · 2026-09-26
- David Sacks: AI regulation lobbying could cost Anthropic and OpenAI their 6-12 month frontier lead — victor_explore · 2026-09-26