A preprint posted to arXiv, bioRxiv, or medRxiv and a peer-reviewed article published from that same underlying research live in two different accountability systems. Journals have a mature, standardized apparatus for handling research misconduct and retraction — investigation triggers, editorial flowcharts from the Committee on Publication Ethics (COPE), and a metadata standard (NISO CREC) that dictates how a retraction notice must be displayed and linked. Preprint servers have nothing equivalent. This guide sets out what actually happens on each side when misconduct is alleged or a retraction occurs downstream, and quantifies the gap between them using the best available empirical study of the problem.
Preprints and journal articles are rarely linked after publication — even when one is retracted
Most researchers assume that if a paper is retracted, the retraction is visible wherever that research appears — including on the preprint that preceded it. In practice, the connection between a preprint and its eventual journal version is maintained inconsistently, and the connection tends to break down completely at exactly the moment it matters most: when the published version is retracted.
This is not a hypothetical concern. It has been measured directly.
The accountability gap, quantified: the Avissar-Whiting study
The most direct evidence on this question comes from a 2022 study published in PLOS ONE, “Downstream retraction of preprinted research in the life and medical sciences,” by Michele Avissar-Whiting. The study matched preprints on three major servers — Research Square, bioRxiv, and medRxiv — to entries in the Retraction Watch Database, covering the period from November 2013 to November 2021. The findings:
- 30 retractions were identified across the three servers: 17 at Research Square, 11 at bioRxiv, and 2 at medRxiv — a very small share (about 0.01%) of all content posted, but the study’s interest was in how those 30 cases were handled, not their raw frequency.
- Only 11 of the 30 (37%) retractions were clearly flagged on the preprint server itself.
- Of those 30, only 5 (17%) had the corresponding preprint separately marked as “withdrawn.”
- On the journal side, only one of the 30 retraction notices visibly acknowledged that a preprint version of the work existed at all.
- The average time from journal publication to retraction was 278 days (range 11–993 days) — notably faster than the roughly 839-day average reported for retractions generally, but that speed advantage does nothing for a reader who encounters the preprint independently and has no way of knowing a retraction happened downstream.
- 70% of the retractions were due, at least in part, to ethical or procedural misconduct rather than honest error, and in 63% of cases the retraction notice indicated the paper’s conclusions were no longer reliable.
The author’s own conclusion is the core of the accountability gap in one sentence: preprint servers and journals each treat the link between a preprint and its published version as the other party’s responsibility to maintain, and in practice neither reliably does. The paper recommends automated linking via Crossref metadata, explicit journal responsibility for maintaining the preprint-to-publication connection, and a reliable mechanism for propagating retraction and correction notices backward to the preprint record.
What “retraction” means on a preprint server — usually, it doesn’t apply
A large part of the gap is definitional. Journals retract articles: a formal, permanent notice replaces or accompanies the original, following an editorial process (often COPE’s own flowcharts) and, increasingly, the display and metadata conventions set out in NISO RP-45-2024 (CREC) — a prepended “RETRACTED:” in the title, a watermarked PDF, and a distinct, linked retraction notice. See How to Write a Retraction Statement and How Long Does a Retraction Take? for how that process works on the journal side.
Preprint servers generally do not have a formal retraction mechanism at all. What they have instead:
- arXiv does not retract or delete papers once they are publicly announced — even at the request of the author or by moderator action. A withdrawal creates a new, top version marked “withdrawn” that states the reason and no longer links directly to the full text, but every prior version remains part of the permanent record and is still individually accessible. arXiv’s moderation and endorsement processes (see arXiv Endorsement System) screen submissions for subject-area fit and format before posting; they are not designed to investigate misconduct allegations after the fact.
- bioRxiv and medRxiv (both operated by openRxiv since the two servers’ operations merged in March 2025) use a similar “Withdrawn” designation rather than a retraction notice, applied inconsistently in practice — per the Avissar-Whiting data above, most bioRxiv/medRxiv preprints behind a later journal retraction were never marked withdrawn at all.
- Neither server operates anything resembling an institutional research-integrity investigation. When misconduct is alleged, the server’s practical options are limited to withdrawing or annotating the preprint; determining whether misconduct actually occurred is not something a preprint server is staffed or positioned to do.
This is a genuinely different function from a retraction statement or an expression of concern as journals use those terms, and conflating “withdrawn” with “retracted” understates how thin the preprint-side mechanism actually is.
Where the gap comes from structurally
Several structural factors compound into the gap the Avissar-Whiting data measures:
- No editorial ownership. A journal retraction is initiated and adjudicated by an editor, often guided by COPE’s decision flowcharts and, for material that first appeared as a preprint, COPE’s own guidance for handling material previously posted to a preprint server. A preprint server has no editor in that sense — screening staff check scope and format, not scientific validity or misconduct.
- No linking mandate. Nothing requires a journal to disclose, in its retraction notice, that a preprint version exists — and per the study above, in practice almost none do. Nothing requires a preprint server to check, on a fixed schedule, whether a linked published version has since been retracted.
- Metadata infrastructure exists but isn’t consistently used end to end. Crossref can carry retraction metadata and NISO CREC standardizes how it should be structured and displayed, but neither is automatically propagated backward to a preprint DOI record unless someone — author, server staff, or journal — does that work manually.
- Institutional findings don’t automatically reach the preprint. If an institution’s research-integrity office (in the US, following the process the ICMJE and federal Office of Research Integrity framework anticipate) concludes misconduct occurred, that finding drives the journal-side retraction. There is no equivalent, standardized channel that also updates the preprint record.
What preprint servers actually do when misconduct is flagged
In practice, preprint servers rely on a narrower set of tools than journals:
- Author-initiated withdrawal — the most common path; the author (or, per the venue’s policy, any listed co-author) requests withdrawal, and the server posts a withdrawn version with a stated reason.
- Moderator or screening-team action — used when a posted preprint is found to violate the server’s own policies (plagiarism, fabricated content, prohibited subject matter). This is a policy-compliance action, not a misconduct adjudication.
- Community and third-party flagging — post-publication commentary platforms (PubPeer is the most widely used) and direct reader reports surface concerns that a server did not catch at screening. Servers generally do not have the investigative capacity to resolve these themselves and typically direct serious allegations to the authors’ institution.
- Deference to the institution and, eventually, the journal. Because preprint servers are not equipped to conduct fabrication/falsification/plagiarism investigations, the substantive resolution of a misconduct allegation almost always happens elsewhere — at the author’s institution or at the journal that eventually reviews the work — and only reaches the preprint record if someone manually updates it.
This is a narrower toolkit than journals have, and it is why the “withdrawn” rate on preprint servers (17% in the measured sample) trails so far behind the retraction rate on the journal side for the same underlying work.
Practical implications for research administrators, authors, and reviewers
- Don’t treat a clean-looking preprint as evidence the underlying work was never retracted. Check both the preprint’s own status page and the Retraction Watch Database for any linked published version before citing or relying on a preprint — see How to Search the Retraction Watch Database.
- Don’t assume the reverse either. A preprint marked “withdrawn” does not always mean misconduct was found — withdrawal also covers author-requested corrections, superseded analyses, and non-misconduct reasons. Read the stated reason rather than inferring intent from the label alone.
- When submitting work that was first posted as a preprint, disclose that clearly at journal submission — see Preprint Server Policies: Avoiding a “Prior Publication” Conflict and How to Submit to bioRxiv/medRxiv — and update the preprint’s own metadata with a published-version link once the journal article is out, which is the single manual step that closes most of the gap described above for a given paper.
- Institutional research-integrity offices should not assume a preprint self-corrects once a related journal article is retracted. If an investigation results in a journal retraction, someone still needs to separately request withdrawal or an update on the preprint record — it will not happen automatically on either arXiv, bioRxiv, or medRxiv.
- Track post-publication discussion, not just formal notices. Post-publication peer review channels often surface concerns well before any formal withdrawal or retraction is posted on either the preprint or the journal side.
Frequently asked questions
Can a preprint be retracted?
Not in the formal sense journals use the term. arXiv, bioRxiv, and medRxiv use a “withdrawn” designation instead: the top version is marked withdrawn with a stated reason, but on arXiv specifically, prior versions remain part of the permanent, publicly accessible record and cannot be fully removed, even by moderator action.
If a paper is retracted by a journal, does the preprint automatically get flagged?
No. Per the Avissar-Whiting 2022 study of 30 downstream retractions across Research Square, bioRxiv, and medRxiv, only 17% of the corresponding preprints were also marked withdrawn. There is no automated mechanism that propagates a journal retraction back to the preprint record; it requires a manual update by the author, the server, or the journal.
Do journals disclose that a retracted article was first posted as a preprint?
Almost never in practice. The same study found only one of 30 retraction notices visibly acknowledged an associated preprint existed, despite most of those preprints being independently discoverable.
Who investigates misconduct allegations against a preprint?
Not the preprint server itself in any substantive sense — servers can withdraw or annotate a preprint for a policy violation, but determining whether fabrication, falsification, or plagiarism actually occurred is an institutional research-integrity function (or, once the work reaches a journal, an editorial one following COPE guidance), not something arXiv, bioRxiv, or medRxiv are staffed to adjudicate.
Is NISO CREC relevant to preprints?
NISO RP-45-2024 (CREC) standardizes how journals should communicate retractions, removals, and expressions of concern in metadata and display. It was written for the journal-publishing context and endorsed by COPE; preprint servers are not bound by it and have not adopted an equivalent standard for withdrawal notices, which is itself part of the accountability gap this guide describes.







