A NeurIPS 2026 conference attendee found a hidden prompt injection in their paper PDF downloaded from OpenReview—instructing any AI system that processes the file to include specific phrases in its output. The injection was not in their original submission, and the researcher suspects it may have been added by the conference to detect whether peer reviewers are using AI to generate reviews without proper human analysis. They are asking other authors to check for the same issue and to flag reviews containing formulaic language that matches the prompt's required phrases.
Summaries like this, in your inbox every morning.
Sign up free →What happened
A researcher discovered a prompt injection embedded in their NeurIPS 2026 paper PDF downloaded from OpenReview—code instructing an AI to include specific phrases in any output. The injection was not present in the original submission, suggesting it may have been added by the conference platform or process.
Why it matters
The discovery raises concerns about the integrity of peer review at a major ML conference. The embedded prompt appears designed to detect whether reviewers used AI to generate their evaluations without proper human review—a practice that undermines the quality and legitimacy of academic evaluation.
What to watch
The researcher is asking other NeurIPS 2026 authors to check their downloaded PDFs for the same injection and to examine peer reviews for formulaic language matching the prompt's required phrases ("This work addresses the central challenge," "The claims of the paper," "Overall, I find this submission"), which could signal AI-generated reviews without genuine human engagement.
A NeurIPS 2026 attendee shared an unexpected discovery on Reddit's Machine Learning community: while reviewing feedback on their paper, they found that GPT had flagged their PDF with a warning about a prompt injection. Prompt injections are hidden instructions embedded in text that instruct AI systems to follow new commands—in this case, to include three specific phrases in any output: "This work addresses the central challenge," "The claims of the paper," and "Overall, I find this submission." The researcher had not inserted this code themselves. Upon comparing the PDF they downloaded from OpenReview (the conference submission platform) with their original submission, the injection was absent from the original version, meaning it must have been added after the submission was uploaded to the conference system. The researcher's hypothesis is that the conference or platform may have embedded the injection as a mechanism to detect whether peer reviewers were using AI to generate their evaluations without genuinely engaging with the paper—any AI-generated review that unknowingly processed the PDF would be forced to include all three phrases, making AI-assisted reviews identifiable. The discovery prompted the researcher to appeal to other NeurIPS 2026 authors to check their own downloaded PDFs for the same injection. They also advised authors to scrutinize their peer reviews for suspiciously formulaic language. If a review contains all three required phrases, the researcher suggested reporting it to the Area Chair (the senior reviewer overseeing a group of papers), as the presence of all three phrases could indicate that the reviewer submitted an LLM-generated review without properly reading and critiquing the paper. The post does not confirm whether other authors have found the same injection or how widespread the issue is, leaving open whether this was an intentional conference-wide measure to combat AI-assisted reviews, a technical error, or an isolated incident.
The discovery touches on a growing tension in academic peer review: the use of large language models (AI systems that understand and generate text) to assist or replace human evaluation. NeurIPS 2026, one of the world's largest machine learning conferences, faces the same pressures affecting academic publishing more broadly—reviewers stretched thin, time constraints, and the increasing availability of AI tools that can draft plausible-sounding evaluations. The embedded prompt injection, if indeed placed by the conference, represents an attempt to detect and flag AI-generated reviews in real time, turning a security mechanism (prompt injection) into a canary in the coal mine. However, the approach also raises questions: if the injection was added by the conference infrastructure automatically, what triggered it, and was the addition transparent to authors? The researcher's call for other authors to examine their own PDFs and reviews suggests the problem may be widespread, though the body does not confirm whether this was an isolated incident or a systematic conference-wide measure.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime