OpenAI self-replicating prompt was misreported, Zvi clarifies it wasn't found in the wild
TheZvi · x · 2026-09-30
David Krueger called out fast-news AI accounts including @TheZvi and @AndrewCurran for multiple misreports of the OpenAI safety story: the original disclosure only said OpenAI "found a self-replicating prompt," not that replication was observed in the wild. TheZvi apologized, saying his initial reaction was too dismissive and the person was communicating something real, and invited criticism of his full writeup.
More from Safety
- Google deploys agentic pre-submit scanning across millions of lines of infra code — 92% precision, 3% false positives — dl_weekly · 2026-09-30
- MongoDB proposes internal AI tool registries for agent tool governance — TheTuringPost · 2026-09-30
- Jailbreak researcher VOID publicly teases "I am going to jailbreak you" — VoidStateKate · 2026-09-30
- Google pays ~100 publishers just 0.1% of ad revenue in AI contribution pilot — gaganghotra_ · 2026-09-30
- UK AISI test shows an autonomous AI cyberattack can cost as little as $1.19 — OmarUFlorez · 2026-09-30
- Bill Gates: if a robot replaces a worker, why doesn't it pay into the pension fund? — _akpiper · 2026-09-30