NeurIPS 2026 author says OpenReview PDF contained a prompt injection
Kwangryeol · reddit · 2026-07-24
A NeurIPS 2026 author says GPT detected a prompt injection inside the PDF downloaded from OpenReview, and that the injected text appears not to have been in the original submission.
The author suspects the injection may have been added by the conference workflow and asks others to check whether they see the same thing in their review copies. They also warn that suspiciously formulaic review wording could indicate LLM-generated reviews that were not properly written by humans.
The post includes the exact injected prompt, which instructs the output to contain three specific phrases, and asks whether anyone else found it in the reviewer version of their paper.
More from Safety
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27
- Gary Marcus Proposes Law Requiring AI Firms to Spend 30% of Budget on Alignment — GaryMarcus · 2026-07-27
- AI coding CLI allegedly uploaded private repos, deleted files and credentials without opt-out — thursdai_pod · 2026-07-27
- Chr Szegedy Discusses Slowing Algorithmic Progress Before RSI — ChrSzegedy · 2026-07-27
- Nature study says AI can simulate human behavior and match experts on experiments — RobbWiller · 2026-07-27
- ExploitGym debate says only 60%–70% of benchmark tasks may be solvable, encouraging cheating — dhadfieldmenell · 2026-07-27