Roughly one in five references in scholarly articles suffers from “reference rot” — the source has either vanished (link rot) or the content at that URL has changed since it was cited (content drift) — and the rate is far higher for citations that point at a bare web page rather than a formally archived or persistently identified source. This guide gives you a decision flow for choosing a durable citation method before you submit, a checklist for archiving URLs you have already cited, and a comparison of the tools available for doing it.
What “link rot” actually means
Link rot is what happens when a hyperlink stops resolving to the resource it originally pointed to — the page is deleted, the site is restructured, a domain lapses, or a paywall goes up where content was once open. It is a subset of the broader problem researchers call reference rot, which also includes content drift: the URL still resolves, but the material at that address has been edited, replaced, or otherwise no longer matches what the citing author actually read and relied on. A citation can fail either way. A reader clicking a dead link gets an error page; a reader clicking a “drifted” link gets a page that quietly no longer supports the claim it was cited for, with no warning that anything changed.
This matters specifically for scholarly citation because bibliographies are meant to be a permanent, verifiable record. A citation to a print journal or a book with an ISBN stays checkable for as long as libraries hold a copy. A bare citation to a live web URL has no such guarantee — the resource it points to exists at the discretion of whoever currently controls that server.
How common is it — the evidence
The most frequently cited empirical study of this problem is Klein et al. (2014), “Scholarly Context Not Found: One in Five Articles Suffers from Reference Rot” (PLOS ONE 9(12), DOI: 10.1371/journal.pone.0115253). The study examined roughly one million references across nearly 400,000 STM (science, technology, medicine) articles published between 1997 and 2012, and found:
- About 1 in 5 articles overall had at least one reference affected by reference rot (link rot, content drift, or both).
- Among articles that specifically cited web-at-large resources (not journal articles with their own DOI), roughly 7 in 10 were affected.
- Failure rates were strongly age-dependent: links in articles from 1997 reached failure rates as high as ~80%, versus roughly 13–22% for articles published around 2012.
The direction of that last finding is the practical takeaway: link rot compounds over time. A URL cited today that looks perfectly stable can fail within a handful of years, and the older a citation gets, the higher the odds it no longer resolves to what was originally cited. Last verified: 2026-08-16.
Decision flow: how should you cite this source?
| What you’re citing | Recommended approach |
|---|---|
| A journal article, dataset, preprint, or other output that has a DOI | Cite the DOI (as a resolvable https://doi.org/10.xxxx/... link), not the URL of the page you happened to land on. See DOI vs. URL for why the DOI is the durable identifier and the URL is not. |
| A web-only resource with no DOI (a government report, a news article, an organization’s policy page, a dataset without a formal identifier) | Archive the URL at the time of citation (Perma.cc or the Internet Archive’s Wayback Machine — see below), and cite both the live URL and the permanent archive link. |
| A resource likely to change content without changing its URL (a wiki page, a “living document,” a policy page a government agency edits in place, a dashboard) | Archive it explicitly and cite the archived snapshot as the source of record, with the access date. The live URL alone will not reproduce what you actually read. |
| Your own institution’s or a funder’s internal/paywalled page | Archive it if your archiving tool can authenticate through the paywall (Perma.cc supports this for many institutional subscribers); otherwise note in the citation that access is restricted. |
| Social media, a forum post, or another inherently ephemeral source | Archive at the time of citation; these are among the highest-risk sources for both link rot and content drift, and platforms routinely delete or edit posts. |
Archiving a cited URL: the checklist
- Archive before you cite, not after. Once a manuscript is submitted, it is too late to capture the page as it existed when you actually consulted it — archive at the moment you first use the source, or immediately before submission at the latest.
- Use a persistent archiving service, not a personal screenshot or PDF save. A local file has no independently verifiable timestamp and no public, checkable URL for readers to follow. A screenshot also cannot be crawled, linked, or resolved the way an archived snapshot can.
- Record the access date. Most citation styles (APA, MLA, Chicago) either require or recommend an access/retrieval date for web sources precisely because the content at a URL is not guaranteed to be stable — check your target journal’s or style guide’s current requirement rather than assuming.
- Cite the archived link alongside, not instead of, the live URL where the citation style allows it, so a reader can see both where the source currently lives and what it looked like when you cited it.
- Re-check archived links before final submission or before a thesis/dissertation deposit — confirm the archive snapshot itself still resolves, since some free archiving tools have historically gone through outages or shutdowns (see WebCite below).
Tools for archiving cited URLs
Perma.cc
Perma.cc is a link-preservation service built by the Harvard Law School Library Innovation Lab, originally to solve link rot in legal citation (a problem documented extensively in US law reviews and court opinions, where a large share of cited web sources go dead within a few years of publication). A user submits a URL; Perma.cc crawls and stores a permanent snapshot, then issues a short, stable perma.cc/XXXX-XXXX link that resolves to that snapshot indefinitely, independent of whether the original page survives. It is now used well beyond law — by academic libraries, journals, and individual researchers — as a general-purpose citation-archiving tool. Access has historically included a free tier for individual registered users with a capped number of links, alongside paid institutional/library subscriptions that raise or remove that cap; check Perma.cc’s own site for current tier limits and pricing before relying on a specific figure, as these terms can change.
Internet Archive Wayback Machine
The Wayback Machine, run by the nonprofit Internet Archive, is a free, general-purpose web archive that has been capturing snapshots of publicly accessible web pages since the mid-1990s, both through its own periodic crawling and through on-demand “Save Page Now” submissions from any user. Unlike Perma.cc, it is not citation-specific and a given page may or may not have been captured at the exact moment you need, but its “Save Page Now” feature lets you force a fresh, timestamped capture on demand, which then resolves at a stable web.archive.org/web/[timestamp]/[original-url] address.
WebCite — no longer a viable option
WebCite (webcitation.org), launched in 1998, was an early on-demand scholarly archiving service and was cited by name in some journals’ author instructions for years. It is not a workable choice today: the service went through an extended outage from October 2021 to June 2023, and as of that point stopped accepting new archive requests altogether — it now only allows viewing snapshots archived before the outage, and remains reported as intermittently unreachable even for that. Any reference relying on a “cite this with WebCite” instruction in an older journal style guide should be redirected to Perma.cc or the Wayback Machine instead. Last verified: 2026-08-16 (via secondary/tertiary sourcing; treat as REPORTED-tier and reconfirm directly against webcitation.org if citing WebCite’s exact operational status again).
Why DOIs solve this more durably than any archiving tool
Archiving services solve link rot after the fact, for a specific snapshot, at a specific URL. A DOI solves a related but distinct problem at the infrastructure level: it is a registered, persistent identifier for a resource that is designed to keep resolving to the resource’s current location even if that location changes, because the registration agency (Crossref for most journal literature, DataCite for datasets and many other outputs) requires the publisher to keep the DOI’s target URL updated whenever content moves. This is handled through the Handle System, the resolution infrastructure underneath every DOI. See DOI vs. URL for the fuller comparison and Crossref vs. DataCite for how the two main registration agencies differ. In practice: if a formal identifier exists for what you’re citing, use it instead of (or in addition to) the URL; if it doesn’t, archiving is the next-best durable option.
Frequently asked questions
What is the difference between link rot and reference rot?
Link rot specifically means a URL no longer resolves at all (404, expired domain, deleted page). Reference rot is the broader term covering both link rot and content drift — cases where the URL still resolves but the content has materially changed since it was cited. Klein et al. (2014) used “reference rot” to capture both failure modes together.
Does using a DOI eliminate the need to archive a source?
For sources that have a genuine DOI issued by a registration agency (Crossref, DataCite, etc.), yes in practice — the DOI is maintained to resolve to the current, correct location for as long as the registration is active, which is the whole point of the infrastructure. Archiving is the recommended fallback specifically for sources that do not have a DOI: web pages, reports, datasets without a formal identifier, and similar.
Is a URL with an access date enough, without archiving it?
An access date tells a reader when you consulted the source, but it does nothing to preserve what the source looked like at that moment — if the page later changes or disappears, the access date alone can’t reproduce it. Archiving (Perma.cc, Wayback Machine “Save Page Now”) captures the actual content; an access date is a useful complement to that, not a substitute for it.
What should I do if a source I already cited has gone dead after publication?
Check whether the Wayback Machine happened to capture it independently at web.archive.org even if you didn’t archive it yourself at citation time — this recovers a surprising number of otherwise-dead citations. If no archived version exists anywhere, note in any errata or corrigendum process your journal supports that the source is no longer accessible, rather than leaving a silently broken citation.







