Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

The Research Data Lifecycle: From Collection Through Archiving and Reuse

A stage-by-stage walk through the research data lifecycle — planning, collection, processing, analysis, sharing, preservation, and reuse — linking to CASRAI’s existing DMP, FAIR, repository, and CRIS content at the point each applies.

Most research data management guidance is organized by document — write a DMP, write a data availability statement, choose a repository. But the underlying data doesn’t move through those documents in isolation; it moves through a sequence of stages, and a decision made at one stage constrains what’s possible at the next. Data collected without a persistent identifier plan is hard to make Findable later. Data processed without a retained script is hard to make Reusable later. This guide walks through that sequence directly — planning, collection, processing, analysis, sharing/publication, preservation/archiving, and reuse — and links to CASRAI’s existing DMP, FAIR, repository, and CRIS content at the point in the lifecycle where each actually applies, rather than restating it here.

Why stages, not documents

CASRAI’s own Data Management Plan (DMP) definition already names the stages explicitly: a DMP describes how research data will be “collected, processed, described, stored, shared, preserved, and (where appropriate) destroyed across the lifecycle of a research project.” That’s the same seven-stage shape this guide uses. The reason to think in stages rather than in required documents (DMP, data availability statement, deposit record) is that the documents are checkpoints, not the whole process — a DMP is written once, near the start, but the commitments it makes have to be executed, stage by stage, all the way through to reuse, or the DMP was just paperwork.

1. Planning: the DMP as the lifecycle’s spine

The DMP is where every later stage gets committed to in advance. A well-formed DMP doesn’t just describe what data will be collected — it specifies, per CASRAI’s own Dictionary, a sharing commitment (which datasets, to whom, under what licence, on what timeline), a preservation commitment (which repository, for how long), a retention period, and — where the data is sensitive — a plan for sensitive-data handling. Funder requirements differ in real, practical ways: the NIH Data Management and Sharing Policy has applied since 25 January 2023, and NSF’s plan was renamed from “Data Management Plan” to “Data Management and Sharing Plan” effective 22 January 2026 under PAPPG 24-1. CASRAI’s NIH vs. NSF Data Management Plans guide covers what actually differs between the two in practice; this guide doesn’t repeat that comparison.

Increasingly, the DMP itself is also machine-readable: a machine-actionable DMP (maDMP) expressed against the RDA DMP Common Standard (v1.2, published November 2025) lets tools like DMPTool or DMPonline exchange the plan’s commitments directly with repositories and institutional CRIS/RIM systems, rather than leaving them as free text a human has to re-key at each later stage. The DMP lifecycle itself — drafting at proposal stage, activation during the funded project, closeout and post-project preservation — runs in parallel with, but is not identical to, the data lifecycle described below: the DMP is a document about the plan; the stages below are what actually happens to the data.

2. Collection: capturing data with reuse already in view

Decisions made at the point of collection are the hardest to retrofit later. Two are worth calling out specifically:

  • Descriptive metadata capture. Recording context (instrument, protocol version, collection date, responsible party, associated project metadata such as funder, award number, and identifiers like ORCID/ROR/RAiD) at the moment of collection is far cheaper than reconstructing it months later at deposit time — and it’s the raw material the FAIR principles (Wilkinson et al., Scientific Data, 2016) actually depend on. “Findable” and “Interoperable” are largely a function of how much structured metadata exists, and metadata created after the fact is reliably thinner than metadata captured in the moment.
  • Governance for sensitive or Indigenous-originated data. Where data involves human subjects or is generated in partnership with or about Indigenous communities, collection-stage decisions about consent, access tiering, and control need to be made before collection starts, not retrofitted at sharing time. The CARE Principles (Collective Benefit, Authority to Control, Responsibility, Ethics — published by the Global Indigenous Data Alliance in September 2019) were developed specifically to complement FAIR with governance and self-determination, not to replace it — Indigenous Data Governance and Indigenous Data Sovereignty are the CASRAI Dictionary terms for the underlying concepts.

3. Processing: from raw capture to analysis-ready data

This is the stage CASRAI’s cluster-level RDM overview treats briefly and this guide treats directly, because it’s where reproducibility is won or lost. Processing — cleaning, transforming, merging, and deriving new files from raw captures — is rarely undone by the time a paper is written, which means the processing steps themselves have to be preserved if anyone (including the original researcher, a year later) is going to be able to verify or extend the work. Two concrete mechanisms matter here:

  • Version control for both code and data. A code repository — typically Git-based, exposing full revision history, branches, and (often) issue tracking — is the standard mechanism for keeping processing/analysis scripts auditable. Where the code itself needs a durable citation independent of any single hosting platform’s uptime, the Software Heritage archive (a non-profit initiative based at Inria) crawls and preserves publicly available source code together with its full version-control history, issuing persistent Software Hash Identifiers (SWHIDs) to archived artefacts.
  • Applying FAIR to the software itself, not just the data. The FAIR4RS Software Citation Principles extend the FAIR Guiding Principles to research software specifically, adapting Findable/Accessible/Interoperable/Reusable to software’s distinctive properties — executability, versioning, dependencies — that don’t map cleanly onto a static dataset. A processing pipeline released without any of this is effectively unreproducible even if the input and output data are both open.

4. Analysis: the stage most audits actually fail at

Analysis inherits every decision made during processing, and adds its own: which statistical or computational environment was used, at what version, with what dependencies pinned. This is the specific ground covered by what’s widely referred to as the reproducibility crisis — the widely reported finding that substantial proportions of published research, particularly in biomedical, psychological, and social sciences, fail to reproduce when re-tested. Funder and publisher openness requirements increasingly extend past the dataset to the analysis itself: the Center for Open Science’s TOP Guidelines include tiered signatory-journal requirements that, at their higher tiers, ask for analysis code and materials, not just data. Recording the analysis environment (dependency manifests, container definitions, or at minimum a documented software-version list) at this stage is what makes the eventual data availability statement (stage 5) actually true rather than aspirational.

5. Sharing and publication: from deposit to data availability statement

Sharing is the stage most researchers associate with “data management,” but by the time it happens, most of the decisions that determine whether it goes well were already made in stages 1–4. Two CASRAI resources cover the mechanics of this stage in depth and aren’t repeated here:

Where the data is sensitive or shared under negotiated conditions rather than openly, a Data Sharing Agreement (DSA) or Data Use Agreement (DUA) governs the terms instead of an open licence, and an embargo can delay open availability to a defined date without abandoning the sharing commitment entirely. On deposit, the resulting dataset landing page — the page a DataCite DOI actually resolves to — is what carries the identifiers, version history, and access conditions forward to every later citation of the data.

6. Preservation and archiving: trust, not just storage

“Archiving” is often treated as synonymous with “uploading somewhere,” but preservation and discovery are separate problems with separate solutions. CASRAI’s How to Choose an Open Data Repository guide covers the generalist-vs-domain-specific-vs-institutional decision and the CRIS-interoperability angle in full; the point specific to this stage is that trustworthiness has to be independently verified, not assumed from a repository’s reputation. A trusted digital repository is one whose governance and technical infrastructure have been assessed against a recognised standard — CoreTrustSeal (16 requirements, formed from a 2017 merger of the Data Seal of Approval and World Data System certifications) is the most commonly cited, alongside the nestor seal and ISO 16363. Metadata quality matters as much as storage durability here: the DataCite metadata schema (currently version 4.7, published March 2026) is what makes a deposit’s title, creators, dates, and identifiers machine-readable across repositories and indexing services — including Google Dataset Search, which CASRAI’s dedicated guide covers directly.

7. Reuse: the stage the whole lifecycle exists to reach

A dataset that’s collected, processed, analysed, shared, and archived but never reused has still only completed half the point of open data. Reuse depends on everything upstream: a dataset without adequate processing documentation can be found and downloaded but not actually understood well enough to reuse correctly; a dataset under an ambiguous or missing licence can be found but not legally reused at all. CRIS/RIM interoperability closes the loop back to institutional reporting — CRIS systems using CERIF and OAI-PMH-style harvesting can pull a dataset record (and, where citation tracking is in place, evidence that it was reused) back into an institution’s own research-information system automatically, rather than requiring manual re-entry every time the data resurfaces in someone else’s work.

How the stages actually interlock

The practical failure mode isn’t usually that one stage is skipped outright — it’s that a decision at one stage silently forecloses an option at a later one. A DMP that commits to open sharing but never specifies a repository leaves the archiving decision to be made under time pressure at submission. Data collected without instrument/protocol metadata can still be deposited, but arrives at the reuse stage effectively undocumented no matter how good the repository is. Treating the DMP’s commitments (sections 1, 5, and 6 above) as binding through processing and analysis — not just at collection and at deposit — is what actually closes that gap.

Frequently asked questions

Is the “research data lifecycle” the same thing as a DMP’s lifecycle?

No, and the distinction is worth keeping straight. The DMP lifecycle is about the plan document itself — drafted at proposal stage, active during the funded project, closed out afterward. The research data lifecycle described in this guide is about what actually happens to the data, which the DMP commits to but doesn’t itself perform.

Where does “processing” end and “analysis” begin?

There’s no universal line, and different disciplines draw it differently — but the practical distinction that matters for reproducibility is whether a step is deterministic and repeatable from raw input (processing: cleaning, format conversion, merging) versus interpretive and dependent on a specific analytical choice (analysis: which model, which statistical test, which threshold). Both need version control and documentation; analysis additionally needs the environment/dependency record described in stage 4.

Does depositing data in a repository automatically satisfy a funder’s DMP sharing commitment?

Not automatically — the DMP’s sharing commitment typically specifies conditions (licence, timing, access tier) that the deposit has to actually match. A generalist repository deposit under a restrictive licence doesn’t satisfy a DMP that committed to open CC-BY sharing, even though something was technically deposited.

Is archived data automatically “reusable” once it’s preserved?

No — preservation (the data continuing to exist and remain accessible) and reusability (the data being adequately documented, licensed, and understandable to someone who wasn’t involved in producing it) are different properties. CoreTrustSeal certification speaks to the former; FAIR’s “Reusable” principle and adequate processing documentation speak to the latter.

Does reuse of a dataset get tracked automatically anywhere?

Only where the infrastructure connects: a DataCite DOI with correctly registered relatedIdentifier links (e.g., a citing publication referencing the dataset’s DOI) can surface in citation-tracking tools and, where CRIS interoperability is in place, in the originating institution’s own research-information system. Absent that metadata linkage, reuse of openly licensed data often goes untracked even when it happens.

For the cluster-level view — DMPs, FAIR and CARE, repository trust, and CRIS integration as a whole, rather than walked through stage by stage — see the Research Data Management (RDM) hub.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →