ICML 2026, held in Seoul in early July, drew a record 23,918 submissions — more than double the prior year’s total. Alongside that surge, the conference’s program committee ran a watermarking scheme designed to catch reviewers who used large language models to write reviews in violation of their declared policy. The result, disclosed in a post on the official ICML blog: 506 unique reviewers were found to have used an LLM despite committing not to, and 497 papers were desk-rejected because 398 of those reviewers were also authors whose own submissions depended on serving as reviewers.
The episode is one of the most concrete enforcement actions taken against undisclosed AI use in peer review to date, and it illustrates both a workable detection method and the scale of the problem it was built to catch.
What ICML found
According to ICML’s own account, reviewers were asked to declare one of two policies before reviewing began: Policy A, no LLM use at all beyond basic spell-checking, or Policy B, LLM assistance permitted for understanding a paper or polishing prose, but not for judging quality or drafting review content. Reviewers who selected Policy A were bound not to route any part of the paper or their review through an LLM.
The watermarking check applied to that Policy A group. It flagged 795 reviews — roughly 1% of all reviews submitted — as showing evidence of LLM involvement, traced to 506 unique reviewers. Because ICML uses a reciprocal reviewing model, in which many reviewers are also authors and a reviewer’s own submission can be tied to their review obligations, 398 of the 506 flagged reviewers were reviewers whose own papers were in the submission pool. Their papers were desk-rejected as a consequence: 497 papers in total, roughly 2% of all submissions, since some flagged reviewers had more than one paper under review.
How the watermark worked
The method, as described on the ICML blog, embedded a hidden instruction inside each submission’s PDF — invisible to a human reader but readable by an LLM that ingests the file’s text layer. The instruction told any LLM processing the document to include two specific phrases, drawn at random for each paper from a purpose-built dictionary of 170,000 candidate phrases, somewhere in its output. If a reviewer’s submitted review text for a paper contained that paper’s assigned phrase pair, it was strong evidence the reviewer had fed the paper (or a review draft built from it) through an LLM rather than reading and reviewing it directly, in breach of a Policy A declaration.
ICML reported that pre-deadline testing found frontier LLMs complied with the embedded instruction and reproduced the watermark phrases at success rates above 80%, and that the overall detection approach carried a reported family-wise error rate of 0.0001 across the review pool. Every flagged case was manually verified by the program committee before any desk-rejection decision was made, rather than acted on by the automated signal alone.
Why the distinction between Policy A and Policy B matters
ICML’s two-tier policy is itself notable for research-integrity purposes: it does not ban LLM assistance in peer review outright. Policy B explicitly tolerates using an LLM to help a reviewer understand unfamiliar material or improve the clarity of their own writing, while drawing a firm line at using an LLM to generate the substance of a scientific judgment — whether a paper is novel, correct, or significant. That is broadly consistent with how funders and publishers have converged on generative AI disclosure norms more generally: assistive use is increasingly tolerated with disclosure, while delegating the evaluative judgment itself to a model is not. The 2026 disclosure landscape for GenAI in scholarly authorship follows a similar assistive-versus-substantive line, and publishers operating under the ICMJE generative-AI policy draw a comparable distinction between permitted assistance and prohibited delegation of judgment.
What made the ICML case different from most prior AI-peer-review disclosures is that it was enforced with a technical detection method built for the purpose, rather than relying solely on reviewer self-disclosure or after-the-fact suspicion. That distinguishes it from a related but separate finding also presented at ICML 2026: a peer-reviewed position paper on “paper laundering” in AI-assisted peer review, which showed that cosmetically rewriting a manuscript with an LLM can inflate the score an AI reviewer assigns it, without any enforcement mechanism attached. The two stories are complementary evidence of the same underlying problem — LLMs distorting peer review, from both the author side and the reviewer side — but they document different studies with different methods and different actors.
Scale: a record submission year strained peer review
The 23,918 submissions ICML 2026 received were more than double the prior year’s total, part of a broader surge across major machine-learning venues as the field’s growth and the accessibility of LLM-assisted writing tools have combined to push submission volumes sharply upward. That volume increase is also the backdrop against which the reviewer-policy violations occurred: a larger, more stretched reviewer pool is both more tempting to route through automation and harder to audit manually, which is part of why ICML built a technical detection mechanism rather than relying on the honor system alone.
What this means for research administrators and integrity offices
For institutions and offices that track research integrity and compliance, the ICML case is a useful reference point for two reasons. First, it demonstrates that undisclosed, policy-violating AI use in peer review is measurable and enforceable at scale — not just a theoretical risk — when a venue is willing to build detection into its submission pipeline rather than relying on reviewer attestations. Second, it reinforces that peer-review integrity policies increasingly need to distinguish assistive from substantive AI use explicitly, the way ICML’s Policy A/Policy B split does, rather than issuing a single blanket rule that reviewers either ignore or interpret inconsistently. Journals, funders, and conferences setting or revising their own reviewer policies can look to this framework as a concrete precedent, including the practical question of how a violation should be adjudicated when the reviewer is also an author in the same submission pool.
Frequently asked questions
How many reviewers did ICML 2026 catch using AI against policy?
ICML’s program committee reported that 506 unique reviewers who had declared “Policy A” (no LLM use) were detected, via the watermarking method, to have used an LLM in preparing a review. Of those, 398 were also authors on the submission pool, and their papers were desk-rejected as a result.
How many papers were desk-rejected, and why does that number differ from the reviewer count?
497 papers were desk-rejected — more than the 398 flagged reviewers — because some of those reviewers had more than one paper under submission, and each affected paper was rejected individually.
How did the watermark actually detect AI-written reviews?
Each submission PDF carried a hidden instruction, invisible to human readers, directing any LLM that processed the document to output two phrases drawn from a 170,000-phrase dictionary unique to that paper. A review that reproduced those exact phrases was strong evidence the reviewer had routed the paper through an LLM rather than reviewing it directly.
Is this the same story as the ICML 2026 “paper laundering” research?
No. Paper laundering is a separate, peer-reviewed position paper presented at ICML 2026 showing that cosmetically rewriting a paper with an LLM can inflate the score an AI reviewer gives it — a vulnerability in AI-generated reviews. The watermark-detection story covered here is ICML’s own program-committee enforcement action against reviewers who used LLMs to write their reviews in violation of policy. Both concern AI distorting peer review, but they are distinct findings from distinct sources.
Sources
Figures and mechanism as reported by ICML’s program committee on the official ICML blog (“On Violations of LLM Review Policies,” blog.icml.cc), with submission-count and conference-timing figures corroborated by independent technology press coverage of ICML 2026 in Seoul.







