A new peer-reviewed position paper accepted at ICML 2026 has put a name to a gaming vulnerability in AI-assisted peer review: “paper laundering” — using a large language model to cosmetically rewrite a manuscript, with no substantive change to its scientific content, specifically to inflate the score an AI reviewer assigns it. The finding is one of the more concrete pieces of evidence yet that automated review systems, as currently built, respond to surface style at least as much as to research substance.
The study
The paper, “Stop Automating Peer Review Without Rigorous Evaluation,” is by Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy (Stanford University and Bocconi University), posted to arXiv (arXiv:2605.03202) in May 2026 and revised in July 2026 ahead of its ICML 2026 presentation in Seoul. The authors tested whether AI-generated peer reviews could be manipulated purely through writing style, independent of a paper’s actual content.
Working from 60 randomly sampled ICLR 2026 papers, they used two “launderer” models (GPT-5.1 and GPT-5.4) to rewrite each paper’s LaTeX source without altering its underlying claims or results, then had three “reviewer” models (GPT-5.1, GPT-5.4, and Claude Sonnet 4.5) score both the original and the laundered version. The laundered versions received a mean score increase of +0.45 on a 1–10 scale, a difference the authors report as statistically significant (Wilcoxon signed-rank test, p<0.001) in nearly every model pairing tested. The laundered papers also became measurably more similar to one another in writing style — a roughly 6.5% increase in pairwise embedding similarity (Cohen’s d = 1.02, p<0.0001) — than the unaltered originals.
The same paper separately documents a “hivemind effect”: AI reviewers tend to agree with each other, and with themselves across papers, more than human reviewers do, reducing the diversity of critical perspective peer review is meant to provide.
What the authors actually claim — and what they don’t
It’s worth being precise about the authors’ own framing, because press and social-media coverage of “paper laundering” has sometimes compressed it into a broader claim than the paper supports. The demonstrated finding is narrow and empirical: AI review scores can be moved by cosmetic rewriting alone, and laundered papers converge stylistically. The authors do not present evidence that AI reviewers, as currently deployed, are systematically suppressing novel or unconventional research relative to safe, incremental work — that would require a different study design than the one they ran.
What the authors do offer is a stated concern, not a finding: if paper laundering became widespread, they write, “scientific writing will converge toward whatever style the AI reviewer rewards, risking an intellectual monoculture and discouraging diverse ways of presenting ideas,” adding that such convergence “could homogenize scientific communication in ways that would disadvantage unconventional but valuable research.” That is a forward-looking risk the researchers are flagging as a reason for caution, framed explicitly as a possibility to guard against — not as something their data shows is already happening.
This distinction matters for anyone citing the study: the gaming vulnerability and the stylistic-homogenization effect are demonstrated; the downstream risk to novel research is the authors’ own reasoned caveat about where that homogenization could lead.
An older, separate critique of peer review generally
The authors’ caveat echoes, but is distinct from, a longer-running critique of human peer review’s relationship to novelty. A frequently cited 2016 study in PNAS by Kevin Boudreau, Eva Guinan, Karim Lakhani, and Christoph Riedl (“Is novel research worth doing? Evidence from peer review at 49 journals”) found that unusually novel proposals tended to receive lower average scores from human reviewers, even as they went on to produce more downstream impact. That body of work is about conventional human review, evaluated over years of grant and journal history — it is not evidence about how AI reviewers behave, and the two shouldn’t be merged into a single claim. The value in noting it here is context: concern about conservatism in peer review predates AI review tools by at least a decade, which is part of why the ICML paper’s homogenization warning has landed as a plausible risk worth watching rather than a novel worry invented from nothing.
Why it matters for scholarly publishing
The practical implication for journals and platforms currently piloting or deploying AI-assisted review is that score-gaming through style manipulation is not hypothetical — it has now been demonstrated under controlled conditions, at a measurable effect size, across multiple current-generation models used as both “authors” and “reviewers.” Publishers relying on AI screening or scoring as any part of an editorial decision are, per this study, relying on a signal that authors can move without improving the underlying work. See CASRAI’s comparison of publisher AI-in-peer-review policies for how different publishers are currently drawing the line on AI involvement in review, and the AI in peer review dictionary entry for the underlying definitions and terminology.
The study adds to a fast-moving evidence base on AI’s role in review; CASRAI covered a separate, unrelated 2026 study on journal-level AI peer-review policy adoption rates in “Study: 83% of High-Impact Journals Have AI Peer-Review Policies.” That piece is about how many journals have adopted formal policies governing AI use in review; this one is about a demonstrated weakness in how AI reviewers score submissions once deployed — the two are complementary but should not be conflated.
What to watch
- Whether journals and conferences that use AI-assisted review add explicit safeguards against stylistic gaming, such as detecting unusual similarity to known “laundering” rewrite patterns.
- Whether follow-up research tests the authors’ homogenization concern directly — that is, whether it measurably affects acceptance of genuinely novel or unconventional work, as distinct from measuring stylistic convergence alone.
- How conference and journal AI-use policies (tracked in CASRAI’s publisher policy comparison above) evolve in response to a documented gaming vector, rather than a theoretical one.
Frequently asked questions
What is “paper laundering” in peer review?
It refers to using an AI model to rewrite a manuscript’s language and presentation — without changing its underlying scientific content or results — specifically to raise the score an AI peer reviewer assigns it. The term comes from the ICML 2026 position paper by Baumann, Pei, Koyejo, and Hovy.
Does the study show AI reviewers favor “safe” research over novel work?
Not directly. The study demonstrates that AI reviewer scores can be manipulated by style alone and that laundered papers become more stylistically similar to each other. The authors separately warn that, if left unaddressed, this could eventually disadvantage unconventional research — but that specific downstream effect on novelty is a stated concern, not something the study’s data measures directly.
Is this the same as journals adopting AI peer-review policies?
No. Policy adoption (how many journals have formal rules about AI use in review) and this gaming vulnerability (how AI reviewers can be manipulated once used) are separate questions, covered in separate studies. See the “related coverage” link above for CASRAI’s coverage of the policy-adoption research.







