A study published in The BMJ in January 2026 applied a machine-learning text classifier to 2,647,471 cancer-research titles and abstracts and flagged 261,245 of them — 9.87% (95% CI 9.83–9.90) — as suspected paper-mill output. It is one of the largest single applications of automated screening to a full disciplinary literature to date, and its authors argue the scale involved is already beyond what manual sleuthing or existing publisher-side tooling can realistically cover on their own.
What the researchers did
The study, “Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study” (BMJ 2026;392:e087581), was led by Baptiste Scancar with co-authors Jennifer A Byrne, David Causeur, and Adrian G Barnett, drawing on IRMAR/CNRS (Rennes), the University of Sydney, and Queensland University of Technology. The team trained a BERT (bidirectional encoder representations from transformers) classifier on 2,202 retracted paper-mill articles identified via the Retraction Watch database, plus 3,094 further papers flagged as suspect by image-integrity experts, split into training (70%), optimization (17.5%), and validation (12.5%) sets. Unlike tools that inspect images, statistics, or reference lists, the classifier works purely on the text of titles and abstracts, learning sentence-level phrasing patterns associated with known paper-mill output rather than checking any single verifiable fact.
On internal validation the model reached 0.91 accuracy, 0.87 sensitivity, and 0.96 specificity; external validation held up at 0.93 accuracy. The researchers then applied the trained classifier to essentially the entire indexed cancer-research literature from 1999 to 2024.
The scale of what it found
Beyond the headline 9.87% figure, the study reports several findings relevant to how the problem is distributed:
- Geographic concentration: 177,907 flagged papers carried a Chinese institutional affiliation, representing an estimated 36% of Chinese cancer-research output in the dataset. Iran followed at roughly 20% of its national output flagged.
- Rising trend: the flagged share rose roughly exponentially between 1999 and 2022, exceeding 15% of annual cancer-research output by the early 2020s.
- Not confined to low-tier venues: the authors note flagged papers appearing in high-impact journals, not only in the marginal titles most discussion of paper mills tends to focus on.
- Citation impact: follow-up reporting in Nature on the same dataset found papers flagged as likely paper-mill output received roughly double the citations of papers not flagged — a reminder that a paper mill’s commercial model depends on its output actually being read and cited, not just published.
How this differs from existing paper-mill screening infrastructure
CASRAI has covered the publisher-side infrastructure built specifically to catch paper mills before or shortly after submission — the STM Integrity Hub, the shared screening service that pools tools like Springer Nature’s donated Geppetto (an internal-consistency and tortured-phrase checker) across member publishers, and the cross-publisher United2Act coalition coordinating a broader response. This BMJ study is a different kind of exercise: it is a retrospective academic research application, not a submission-time production tool, and it was run once across an entire discipline’s literature rather than continuously at the point of manuscript intake. Where Geppetto-style tools and manual sleuthing (the kind that populates the Retraction Watch database and drives individual expressions of concern) work case by case, this classifier was applied to 2.6 million records at once — a scale the authors frame as illustrating the size of the detection gap rather than as a deployable screening product in its own right.
It also sits in a different methodological family from statistics-based fraud-detection techniques such as GRIM, SPRITE, or statcheck, which test whether reported numbers (means, standard deviations, p-values) are internally consistent. The BMJ classifier makes no claim about whether any specific number in a paper is correct; it learns phrasing patterns statistically associated with a training set of already-identified paper-mill papers, which is a fundamentally probabilistic, pattern-matching signal rather than a factual check.
Limitations the authors themselves flag
The paper is unusually candid about what the classifier cannot tell you:
- The training data is skewed toward retractions from 2013–2023 and, within that, toward Chinese-affiliated papers, creating a residual risk that the model is partly learning to recognize country-of-origin signals rather than fraud signals specifically.
- The authors estimate roughly 30% of flagged papers may be false positives.
- As a deep-learning model, the classifier is not fully explainable — the researchers cannot point to exactly which textual features drove any individual flag.
- A flag is a probabilistic screening signal, not a misconduct finding; the authors position the tool as triage for human follow-up, not a replacement for investigation.
What it means for publishers and research-integrity teams
For journal editors, publisher integrity offices, and institutional research-integrity teams, the practical takeaway is less “here is a new tool to deploy” and more a scale argument: if a single retrospective academic exercise can flag a quarter-million suspect papers in one discipline, the volume of material that would need triage under any large-scale automated screen is going to substantially exceed current human review capacity, reinforcing the case publishers have already made for pooled infrastructure like the STM Integrity Hub over each title building its own bespoke tooling. It also strengthens the argument, already made by COPE’s relaunched paper mills working group, that sentence-level and AI-generation-adjacent detection needs to be treated as a distinct, fast-evolving front separate from image-manipulation or citation-based checks — see also CASRAI’s coverage of a related peer-review collusion network uncovered through more conventional editorial investigation. With a reported 30% false-positive rate, any institution or publisher considering similar classifiers should plan for confirmatory human review as a mandatory downstream step, not an optional one.
Frequently asked questions
Is this the same technology as STM Integrity Hub or Geppetto?
No. Geppetto and the STM Integrity Hub are publisher-operated, submission-time screening infrastructure. The BMJ study is an independent academic research project that applied a purpose-built BERT text classifier retrospectively across an entire indexed literature; the authors do not present it as production screening software.
Does a “flag” from this model mean a paper is fraudulent?
No. The authors describe it as a probabilistic screening signal with an estimated 30% false-positive rate, intended to prioritize papers for human review, not as a standalone determination of research misconduct.
Where can I read the study itself?
The full paper, “Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study,” is published in The BMJ (2026;392:e087581) and available via PubMed Central.







