Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Data Rescue Projects: Recovering At-Risk Scientific Data Before It’s Lost

What data rescue projects are, why at-risk research and environmental data gets lost, and how real initiatives like Data Refuge, DataRescue, and IEDRO have recovered it.

A data rescue project is a coordinated effort to identify, recover, and preserve research or scientific data that is at genuine risk of being permanently lost—because it exists only on decaying physical media, in an obsolete file format, on a server about to be decommissioned, or on a government website whose continued maintenance is not guaranteed. Data rescue is distinct from routine archiving: routine archiving happens as part of a planned workflow while data is still accessible and its custodian is still engaged; data rescue happens after that window has already closed or is closing, when the data’s survival depends on someone actively intervening before the medium fails, the format becomes unreadable, the institution loses funding, or the hosting page is taken down.

The term covers a wide range of activity: digitizing paper strip-chart weather records before the paper degrades, migrating datasets off obsolete tape formats or unsupported software, downloading and mirroring government datasets ahead of anticipated policy or administration changes, and recovering legacy observational records that were never digitized at all. What unifies these efforts is the same underlying risk: primary research data, particularly long time-series environmental and observational records, cannot be regenerated once lost. A gap in a 100-year climate record cannot be re-measured retroactively; a discontinued federal dataset with no equivalent replacement leaves a permanent hole in the scientific record.

Why data rescue is a distinct problem from ordinary data management

Most research data management guidance, including CASRAI’s own research data lifecycle content, assumes data enters a managed pipeline close to the point of creation: it is described with metadata, deposited in a repository, and preserved under an explicit preservation commitment. Data rescue exists precisely because a meaningful share of the world’s scientific data predates that kind of managed pipeline, or fell outside it—paper logbooks from decades before digital recordkeeping existed, datasets produced under short-lived grants with no ongoing preservation obligation, or agency data published to the web with no long-term archiving mandate attached.

Three risk categories drive most data rescue work:

  • Media decay. Magnetic tape, punch cards, microfiche, and paper records degrade physically. Analog weather and hydrological records held in national meteorological services, especially in lower-resource countries, are a frequently cited example: the paper itself becomes brittle or the ink fades before anyone digitizes the readings on it.
  • Format and hardware obsolescence. Data stored in a proprietary format tied to discontinued software, or on a physical medium (a specific tape drive, an old disk format) for which working read hardware is disappearing, becomes inaccessible even though the bits themselves may still exist. This is the operational problem the OAIS Reference Model (ISO 14721) and standards like PREMIS are designed to prevent going forward, but neither helps with data that was never brought into a managed preservation environment in the first place.
  • Institutional and funding disruption. A dataset hosted on a single agency server, dependent on a single grant-funded maintainer, or published without a mirrored, versioned, independently hosted copy is at risk the moment that hosting arrangement changes—whether through a funding lapse, a change in agency priorities, a website redesign that drops old content, or a change in government policy toward a specific data domain.

Why it matters

The case for data rescue rests on a simple asymmetry: collecting new data is often expensive and sometimes impossible, while losing existing data is often silent and irreversible. A few concrete stakes:

  • Long time-series records cannot be backfilled. Climate, hydrological, and seismic research depends on continuous observational records stretching back decades or centuries. A gap in a historical weather record is a permanent gap—there is no way to re-observe 1950s rainfall in 2026. The World Meteorological Organization has published best-practice guidance specifically because historical instrumental records are considered irreplaceable inputs to climate modeling and trend detection.
  • Reproducibility and scientific continuity. Data underlying published findings that later becomes inaccessible undermines the ability of other researchers to verify, reanalyze, or build on that work—the same reproducibility concern that motivates funder data-sharing and retention requirements more broadly.
  • The cost asymmetry compounds over time. Digitizing or migrating a dataset while the original medium and metadata (or the people who understood the collection context) still exist is far cheaper and more accurate than attempting recovery after context is lost. Delay narrows what’s actually recoverable.
  • Public-good data is not automatically self-sustaining. Much at-risk data is government-collected environmental, meteorological, or public-health data with no commercial owner and no guaranteed long-term maintenance mandate; its survival depends on somebody—an agency, a library, a volunteer network—treating preservation as an active responsibility rather than assuming the data will simply persist.

Notable real-world data rescue initiatives

The following are documented, verifiable examples of data rescue activity, spanning different risk types and different organizational models—grassroots/volunteer, nonprofit/international, and standards-body-coordinated.

DataRescue and Data Refuge (2016–2017)

Data Refuge was launched in November 2016 by the Penn Program in Environmental Humanities at the University of Pennsylvania, in partnership with the Environmental Data & Governance Initiative (EDGI), in response to concern that federal environmental and climate datasets hosted on U.S. government websites could be altered, deprioritized, or taken offline amid an anticipated change in federal environmental policy. Between December 2016 and June 2017, local organizers ran 49 volunteer “DataRescue” events in cities across the U.S. and Canada. Using a browser extension built for the effort, participants nominated more than 63,000 federal web pages as candidates (“seeds”) for web archiving, and identified more than 22,000 datasets as candidates for more hands-on, non-automated preservation; several hundred of those datasets were fully “harvested” through a workflow EDGI and Data Refuge developed and uploaded into the Data Refuge repository. The effort worked alongside related volunteer initiatives including Climate Mirror and #ProtectClimateData. It is a useful case study specifically because the risk driving it was institutional/policy risk rather than physical media decay—the data was born digital and technically accessible right up until it wasn’t.

IEDRO (International Environmental Data Rescue Organization)

IEDRO is a U.S.-based nonprofit that recovers and digitizes historical environmental and weather records—paper strip charts, logbooks, and station records held by national meteorological and hydrological services, predominantly in developing countries, that are physically deteriorating and were never digitized. IEDRO has worked on data recovery with national meteorological services in multiple African and Latin American countries, and coordinates with the World Meteorological Organization (WMO) and with NOAA’s National Centers for Environmental Information, where rescued data is deposited so it becomes freely and permanently accessible. Its work is a clear example of the media-decay risk category: the underlying observations exist nowhere except on physically fragile paper, so the sole recovery window is before that paper is lost.

World Meteorological Organization data rescue guidance

Because climate science depends on long, continuous instrumental records, the WMO and partners including the Copernicus Climate Change Service have published best-practice guidelines specifically for climate data rescue—covering how to locate, image, key (digitize), and quality-control historical records so that recovered data is actually usable in downstream climate analysis, not just preserved as a scanned image. This standards-body involvement reflects that data rescue, done well, is a methodological discipline (with its own quality-control and metadata practices) and not simply a matter of copying files before they’re lost.

What a data rescue effort typically involves

Across these different initiatives, a recognizable set of steps recurs:

  1. Risk identification. Someone has to notice that a specific dataset, collection, or record series is at risk—whether because of physical condition, an impending institutional change, or format obsolescence—before the loss actually occurs. This is often the hardest step, since at-risk data is by definition not being actively monitored by anyone with a preservation mandate.
  2. Triage and prioritization. Rescue capacity is always finite relative to the volume of at-risk material, so efforts typically prioritize by irreplaceability (is there any other copy anywhere?), scientific or public value, and how imminent the loss risk is.
  3. Capture. Depending on the risk type, this means web archiving/mirroring (for born-digital, institutionally-hosted data), imaging or scanning (for paper records), or format migration and bit-level copying (for obsolete digital media).
  4. Digitization and keying. For non-digital sources, an imaged record still has to be transcribed or algorithmically extracted into structured, machine-readable data—the step that turns a scanned strip chart into an analyzable time series.
  5. Quality control and metadata. Rescued data needs documentation of its provenance, collection method, and any transcription or migration steps applied, or it risks becoming unusable or untrustworthy even once technically recovered—the same concern the OAIS model addresses for planned preservation workflows.
  6. Deposit in a durable, independently-governed home. Rescued data is generally deposited somewhere with a longer-term preservation mandate than the original at-risk location had—a national data center, a trusted digital repository, or a certified archive—specifically so it doesn’t become at-risk again in the same way.

How data rescue relates to routine research data management

Data rescue and planned RDM sit on the same continuum but at opposite ends of it. A well-executed data management plan with a genuine preservation commitment, deposit in a CoreTrustSeal-certified repository, and routine fixity checking is, functionally, data-rescue prevention: it is designed so that the dataset never enters the at-risk state that rescue projects exist to respond to. For research administrators and data stewards, the practical takeaway is that data rescue is what happens when that upstream planning didn’t occur, wasn’t sufficient, or the world changed in ways the original plan didn’t anticipate (an agency reorganized, funding lapsed, a policy shifted). Institutions and funders increasingly treat evidence of a credible long-term preservation plan—not just short-term storage—as part of what makes a repository choice and a DMP adequate in the first place, which is the direct, preventive counterpart to the reactive work described on this page.

Frequently asked questions

Is data rescue the same as digital preservation?

They’re related but not identical. Digital preservation, as formalized in standards like the OAIS Reference Model, is the ongoing set of practices that keep already-managed digital data accessible and usable over time. Data rescue is narrower and more urgent: it specifically addresses data that fell outside, or predates, any such managed preservation process and is at active risk of being lost before it can be brought into one.

Who typically carries out data rescue projects?

It varies by risk type and scale: volunteer and academic networks (as with DataRescue/Data Refuge, coordinated through universities and civil-society organizations), specialized nonprofits (as with IEDRO’s work on physical meteorological records), and international standards bodies and national agencies (as with WMO-coordinated climate data rescue guidance and NOAA’s role receiving rescued climate records). Larger, ongoing efforts often involve some combination of all three.

What kinds of data are most often the subject of rescue efforts?

Long environmental and observational time series—weather, climate, hydrological, and oceanographic records—are especially common subjects, both because they are scientifically valuable specifically for their length and continuity, and because a meaningful share of the historical record predates digital recordkeeping and exists only on paper or obsolete media. Government-hosted datasets facing institutional or policy-driven discontinuation are another recurring category, illustrated by the scale of the 2016–2017 DataRescue/Data Refuge events.

How does a research institution avoid needing a future data rescue effort for its own data?

By treating the steps that make rescue unnecessary as part of routine research data management rather than as optional extras: depositing data in a certified, independently governed repository rather than only on a lab’s own server; documenting preservation commitments explicitly in a data management plan; and avoiding single points of institutional failure (one grant, one server, one person) as the sole custodian of a dataset with lasting value.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →