Researcher corrects viral claims: OpenAI's self-replicating prompt was found in academia 2 years ago
DavidSKrueger · x · 2026-09-30
David Krueger (Cambridge) flags two misreports by popular AI news accounts about OpenAI's self-replicating prompt story:
- OpenAI's note only says "we found a self-replicating prompt" — it was NOT found replicating in the wild (though Krueger wouldn't be surprised if that's happening undocumented).
- It's not even a novel OpenAI discovery: academics found this two years ago in "Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems" (Lee & Tiwari).
Related event: Cambridge Researcher Clarifies Prompt Self-Replication Was Studied in 2024(2 posts)→
More from Safety
- Berkeley issues AI-use disclosure guidance down to paper sections and figures — hoofnagle · 2026-09-30
- PromptBrake launches free MCP server for prompt injection payloads and OWASP LLM risk mapping — Specialist-Bee9801 · 2026-09-30
- Study of 100K crime-forum posts: AI speeds up cheap scams, doesn't create new hackers — rohanpaul_ai · 2026-09-30
- AI agents leak 13,000+ internal screenshots from 343 tech companies to public GitHub repos — AccBalanced · 2026-09-30
- Anthropic, Meta and Google sign voluntary White House AI safety development pact — dotey · 2026-09-30
- Musk details joint AI safety declaration with cross-company monitoring and peer review — XFreeze · 2026-09-30