A Lancet Correspondence published May 7, 2026 puts a hard number on a trend research-integrity investigators have been flagging anecdotally for two years: fabricated citations in the biomedical literature are rising sharply, and the timing tracks almost exactly with the mainstream adoption of large language models for manuscript drafting.
What the audit found
A team led by Maxim Topaz of Columbia University’s Data Science Institute and School of Nursing audited roughly 2.5 million PubMed Central open-access papers published between January 2023 and February 2026. Using an automated checker that cross-referenced citation details against PubMed, Crossref, OpenAlex, and Google Scholar, the team verified 97.1 million individual references and flagged any that could not be matched in any of the four databases as fabricated — as distinct from references that were merely mis-formatted or used informally abbreviated journal titles.
The result: 4,046 fabricated references across 2,810 papers. The paper-level fabrication rate rose steadily across the study window:
- 2023: approximately 1 in 2,828 papers contained at least one fabricated citation.
- 2025: the rate climbed to roughly 1 in 458 papers.
- Early 2026: the rate reached 1 in 277 papers — more than a tenfold increase since 2023.
The authors identified mid-2024 as the inflection point, which coincides with the point at which generative AI writing tools moved from novelty to routine use in manuscript preparation. Several secondary findings sharpen the picture: 91% of affected papers contained only one or two fabricated references rather than a large volume, suggesting most instances are a hallucinated citation slipping through alongside otherwise legitimate ones rather than wholesale fabrication. Review articles — which typically cite far more sources than original research — showed a fabrication rate 57% higher than other article types. And fabricated references were disproportionately concentrated in large open-access journals, which publish at higher volume and, in some cases, operate leaner editorial screening per submission.
Perhaps the most consequential finding for editorial policy: as of the audit’s cutoff, over 98% of the flagged papers had seen no publisher action — no correction, no expression of concern, no retraction — despite the fabricated references being identifiable through the same automated cross-referencing the researchers used.
Why LLMs fabricate citations specifically
The mechanism behind this pattern is well understood in the AI research community, not a mystery unique to this dataset. Large language models generate text by predicting statistically plausible continuations, not by retrieving verified facts from a database. A citation to a specific author, journal, volume, and page range is, to the model, just another plausible-sounding string — it can be stylistically indistinguishable from a real citation while referring to nothing that exists. Researchers have called this fabricated or “hallucinated” citation phenomenon a structural feature of how current LLMs operate, not an occasional bug — a NeurIPS 2025 paper on compound deception in elite peer review (‘Compound Deception in Elite Peer Review,’ arXiv:2602.05930) proposed a taxonomy of exactly this class of fabricated-citation and fabricated-review-artifact behavior, distinguishing it from more overt forms of fraud precisely because the outputs are designed, by the nature of the generation process, to look legitimate.
That is also why fabricated citations are comparatively hard to catch by traditional peer review. A reviewer skimming a reference list for plausibility — correct-looking author names, a real journal title, a year that fits the topic — has no reliable way to tell a hallucinated entry from a real one without individually looking each one up. That is precisely the gap the automated cross-referencing approach in the Lancet audit, and the citation-screening tools now being adopted by some publishers, is built to close.
How publishers and platforms are responding
Response has been uneven and, per the audit’s own finding, largely absent at the individual-paper level so far. Some publishers have been more concrete than others. Frontiers, for instance, has had an in-house AI integrity screening system called AIRA (AI Research Assistant) in production since 2018, running dozens of automated checks against every submission before it reaches peer review — plagiarism, image manipulation, papermill signals, and increasingly citation and reference-pattern anomalies — with a human editor always making the final call on any flag. In 2025 Frontiers expanded AIRA further, integrating third-party fraud-screening tools (Cactus Communications’ Paperpal Preflight and Clear Skies’ Papermill Alarm and Oversight) and announcing enhanced reference and citation checks specifically aimed at this category of problem. PLOS has said it is exploring system-wide reference-integrity screening, while also cautioning — correctly, from a due-process standpoint — that a fabricated reference in a manuscript is not automatically equivalent to a finding of research misconduct against the author; intent and degree still matter under standard definitions like fabrication under 42 CFR Part 93.
For an editorial office, the practical takeaway is that citation-count spot-checks by reviewers are no longer sufficient at current fabrication rates, and that the same kind of automated database cross-referencing used in the Lancet audit is realistically the only scalable way to catch this at submission. Whether that becomes a submission-gate requirement industry-wide, the way AI-text detection and image-integrity screening increasingly are, is the open question the next 12-18 months of publisher policy will answer.
What this means for authors, editors, and institutions
For individual authors using AI writing assistants, the practical risk is concrete: an LLM-drafted paragraph can insert a citation that reads perfectly and does not exist, and unless every reference is independently verified against a real database before submission, that fabricated citation ships under the corresponding author’s name. See CASRAI’s guide on citation management with AI writing tools for verification workflow guidance, and the related walkthrough on how AI-generated text is treated under plagiarism and misconduct policy.
For editors and integrity officers, the 98%-no-action figure is arguably the more urgent finding than the headline rate itself: detection tooling now exists and is affordable to run at scale, but institutional and editorial response has not caught up to deployment. That gap — between what can be detected and what is actually being corrected — is likely to be where policy attention concentrates through the rest of 2026.
Frequently asked questions
Is a fabricated citation the same as plagiarism?
No. Plagiarism is presenting someone else’s work or words as your own; a fabricated citation is a reference to a source that does not exist at all, typically generated by an LLM predicting a plausible-looking author/title/journal string rather than retrieving a real one. The two can co-occur but are distinct problems with distinct detection methods.
Does a fabricated citation automatically count as research misconduct?
Not automatically. Under frameworks such as 42 CFR Part 93’s definition of fabrication, misconduct findings generally require intent, and publishers including PLOS have explicitly noted that a hallucinated reference introduced via an AI writing tool is not the same, on its face, as deliberately fabricating data. Institutions and journals still need to assess the individual case — how the citation entered the manuscript, and whether the author took reasonable steps to verify their references before submission.
How can authors avoid submitting AI-hallucinated citations?
Every citation generated or suggested by an AI writing tool should be independently verified against a bibliographic database (PubMed, Crossref, OpenAlex, or the publisher’s own indexing) before submission, the same cross-referencing approach the Lancet audit itself used. Relying on an LLM’s own confidence that a citation is real is not a verification step, since the model has no mechanism for knowing the difference between a real and a fabricated reference.
Which journals or fields are most affected?
The audit found fabricated references concentrated disproportionately in large open-access journals and in review articles specifically, which showed a fabrication rate 57% higher than other article types — plausibly because review articles cite far more sources per manuscript, giving more opportunities for a hallucinated reference to appear undetected.







