NeurIPS 2026 author says OpenReview PDF contained a prompt injection
Kwangryeol · reddit · 2026-07-24
A NeurIPS 2026 author says GPT detected a prompt injection inside the PDF downloaded from OpenReview, and that the injected text appears not to have been in the original submission.
The author suspects the injection may have been added by the conference workflow and asks others to check whether they see the same thing in their review copies. They also warn that suspiciously formulaic review wording could indicate LLM-generated reviews that were not properly written by humans.
The post includes the exact injected prompt, which instructs the output to contain three specific phrases, and asks whether anyone else found it in the reviewer version of their paper.
More from Safety
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11
- Fields Medalist founds Mathematical AI Safety Institute to prove AI safe like cryptography — The Decoder · 2026-09-11