Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Research Data Governance: Roles, Policies, and Lifecycle Frameworks

What research data governance means, the roles it assigns (owner, steward, custodian, user), the policy layer it sits on top of, and the lifecycle frameworks used to operationalize it.

Research data governance is the framework of decision rights, roles, and policies that determines who is accountable for a dataset, who may access or modify it, and under what conditions it can be shared, retained, or destroyed. It is a distinct layer from data management: management is the day-to-day operational work of organizing, documenting, and storing data (the work a Data Management Plan (DMP) commits a project to); governance is the accountability structure that decides who has authority to make those calls, and enforces that they’re followed. A project can have excellent data management practice and still have no governance — no one formally accountable, no documented access policy, no answer to “who decides” when a dispute arises — and the reverse is just as possible. This guide covers the general, institution-level concept: the roles governance frameworks typically assign, what a governance policy actually specifies, how access control fits in, and the lifecycle models used to structure the whole thing. It links out to CASRAI’s more specific existing content — the data lifecycle, the data steward/curator role, data ownership, and the CARE Principles for Indigenous data governance — at the point each becomes relevant, rather than repeating it here.

Governance vs. management vs. stewardship

These three terms get used loosely and interchangeably in practice, but they describe different layers of the same system:

  • Governance is the decision-rights and accountability layer: who is authorized to set policy, approve access, and answer for a dataset if something goes wrong. It’s typically expressed as a policy document, a committee or designated role, and an enforcement mechanism.
  • Management is the operational execution of that policy: the concrete work of collecting, describing, storing, and sharing data according to a plan. See CASRAI’s research data lifecycle guide for how that work breaks down stage by stage.
  • Stewardship sits between the two — an ongoing, hands-on accountability for a dataset’s quality, documentation, and accessibility across its life, usually carried out by a named data steward or curator role. CASRAI’s data stewardship and data curator role guide covers this specific job function in depth, including how “steward,” “curator,” and “manager” job titles differ across institutions (there’s no single universal distinction — institutions that do distinguish deliberately tend to use “curator” for hands-on metadata/format work near deposit and “steward” for broader lifecycle accountability).

Governance frameworks in the broader data-management field — DAMA International’s Data Management Body of Knowledge (DAMA-DMBOK) is the most widely cited — define data governance in essentially this same way: the exercise of authority and decision-making over the management of data assets, distinct from the operational management work itself. Research data governance applies that same structure to a narrower, higher-stakes domain: data generated by, or in the course of, research, often subject to funder mandates, human-subjects protections, and institutional ownership policy that general enterprise data doesn’t carry.

Core roles: owner, steward, custodian, user

A functioning governance framework assigns each of these roles explicitly, even if one person holds more than one at a small scale:

  • Data owner — the party with ultimate authority to set policy for a dataset: who may access it, under what license it may be shared, and how long it’s retained. Ownership of research data is genuinely unsettled at the level of general law: CASRAI’s data ownership entry covers this, but in short, US federal regulation (2 CFR §200.315, the OMB Uniform Guidance) governs *rights in* data produced under a federal award — the recipient may copyright the work while the federal awarding agency retains a royalty-free right to use it for federal purposes — without itself assigning default ownership of the underlying data to either party. The NIH Data Management and Sharing Policy (NOT-OD-21-013) likewise governs planning and sharing, not ownership. In practice, ownership is set by institutional policy: several US research universities, including the University of California (Research Data Policy, effective July 2022) claim institutional ownership of research data created by their researchers in the course of university research, with the principal investigator treated as the data’s primary steward rather than its legal owner — a deliberate distinction between who controls day-to-day decisions and who holds ultimate title.
  • Data steward — the role with ongoing, working-level responsibility for a dataset’s quality, documentation, and accessibility: maintaining the DMP, liaising on ethics/consent questions, and making sure metadata and access conditions stay current across the project’s life. See CASRAI’s data steward role in a DMP entry for how this is typically written into a plan.
  • Data custodian — the role (often IT, a repository, or a research computing unit) responsible for the technical implementation of governance decisions: applying the access controls the owner/steward specify, maintaining backups and security, and executing retention/disposal schedules. The custodian implements policy; it does not usually set it.
  • Data user — a researcher, collaborator, or third party granted access under the governance policy’s terms, typically bound by a data use agreement or data sharing agreement when access crosses institutional or organizational lines.

Note that HHS’s Office of Research Integrity (ORI) treats “ownership,” “control,” and “rights” largely interchangeably in its own guidance and doesn’t formally separate ownership from stewardship/custodianship as distinct concepts the way this guide does — the owner/steward/custodian/user breakdown above is a widely used practical model in the research-data-governance literature, not a single universally codified legal taxonomy. Confirm how a specific institution or funder actually defines these roles before assuming this exact structure applies.

What a research data governance policy actually covers

At an institutional level, a research data governance policy typically specifies:

  • Role assignment — who is the default owner, steward, and custodian for data generated within the institution (often by default: the institution owns, the PI stewards, a central IT/library function custodianships the infrastructure), and how that changes for externally funded, multi-institution, or human-subjects data.
  • Classification and sensitivity tiers — categorizing data by sensitivity (e.g., open, restricted/controlled-access, confidential/identifiable) so that access rules can be applied consistently rather than negotiated case by case. Controlled-access biomedical repositories such as dbGaP (NIH’s Database of Genotypes and Phenotypes) illustrate this in practice: open/unrestricted tier data (study documentation, summary statistics) requires no request, while individual-level genotype/phenotype data sits in a controlled-access tier requiring Data Access Committee approval.
  • Access control — the mechanism by which the classification tiers above are actually enforced: role-based permissions, data access committees, or agreement-gated release. Where data concerns identifiable individuals, this overlaps with privacy frameworks such as the US Fair Information Practice Principles (FIPPs) and, in the EU/UK, statutory research exemptions and safeguards under data protection law.
  • Retention and disposition — how long data must be kept (often set by funder or institutional policy, commonly several years post-publication or post-award) and the process for secure disposal once that period lapses.
  • Sharing and reuse conditions — what license or agreement governs external reuse, and who is authorized to approve a sharing request. This is where governance policy connects directly to a project’s DMP: the DMP is where these commitments get written down for a specific award; governance policy is what makes those commitments consistent across all of an institution’s projects rather than reinvented per grant.

Lifecycle governance: applying policy across the data’s life

Governance isn’t a single decision made at deposit; a dataset’s risk profile and appropriate access rules can change at every stage of its life. The Digital Curation Centre’s (DCC) Curation Lifecycle Model — conceptualize, create/receive, appraise and select, ingest, preservation action, store, and access/use/reuse — is the framework most widely used in the research-data field to structure this, and CASRAI’s own research data lifecycle guide walks through it stage by stage. Governance-relevant decisions recur at several of these stages specifically: appraisal (should this dataset be retained at all, and at what sensitivity tier), ingest (does the receiving repository’s access-control model match the governance policy’s requirements), and access/reuse (has the licensing and agreement structure been correctly applied before release). A repository’s own governance maturity is itself assessed as part of trusted-repository certification: CoreTrustSeal’s sixteen requirements include a dedicated “Governance and Resources” requirement (R05) alongside “Rights Management” (R02) and “Legal and Ethical Compliance” (R04) — see CASRAI’s CoreTrustSeal entry — specifically because a repository can’t be trusted with governed data unless its own internal governance is sound.

Where research data governance intersects specialized frameworks

The general framework above is deliberately broad; several more specific governance frameworks apply in particular contexts, and this page defers to them rather than duplicating them:

  • Indigenous data — data about, or from, Indigenous peoples is subject to a distinct governance framework, the CARE Principles for Indigenous Data Governance (Collective Benefit, Authority to Control, Responsibility, Ethics), developed by the Global Indigenous Data Alliance (GIDA). CARE is explicitly complementary to, not a replacement for, the FAIR principles: FAIR addresses data’s technical findability/accessibility/interoperability/reusability, while CARE addresses who has the authority to govern it. See CASRAI’s Indigenous Data Governance and FNIGC entries for the fuller treatment.
  • AI training data — the EU AI Act imposes its own, narrower data-governance obligation under Article 10, requiring specific data governance and quality-management practices for the training, validation, and testing datasets used in high-risk AI systems. That’s a governance requirement for a specific regulated purpose (AI system compliance), not a general research-data governance framework — see CASRAI’s EU AI Act Article 10 entry if that’s the specific obligation you’re looking for.
  • Funder-specific DMP requirements — NIH and NSF, among others, set distinct requirements for what a data management and sharing plan must address; see CASRAI’s NIH vs. NSF data management plans comparison for how those differ in practice.

Building a governance framework in practice

An institution setting up (or auditing) a research data governance framework is typically working through a version of this sequence: (1) inventory what data the institution generates and classify it by sensitivity; (2) assign the owner/steward/custodian roles above, by default and by exception (e.g., multi-institution collaborations, industry-sponsored research); (3) write the access-control and retention rules that follow from that classification, ideally as a standing policy rather than negotiated per project; (4) establish an oversight mechanism — a data governance committee, or an existing body such as research integrity or IT security taking on the function — with actual authority to resolve disputes and approve exceptions; and (5) connect the policy to the DMP process, so that a project’s DMP is checked for consistency with institutional governance policy rather than existing as a separate, unenforced document. None of this replaces a project-level DMP; it’s the layer that makes DMPs consistent and enforceable across an institution rather than a one-off promise per award.

Frequently asked questions

What’s the difference between data governance and a data management plan?

A DMP is a project-level document describing how one project’s data will be handled. Data governance is the institution-level (or funder-level, or repository-level) policy and role structure that a DMP has to be consistent with. A DMP is written once per award; governance policy applies across all of an institution’s research data continuously.

Who owns research data — the researcher or the institution?

There’s no single universal answer; it depends on institutional policy and, where relevant, the funding award’s terms. US federal Uniform Guidance (2 CFR §200.315) governs certain rights in federally funded research outputs but doesn’t itself assign ownership of the underlying data. Many US research universities claim institutional ownership by policy, with the principal investigator as steward rather than owner. See CASRAI’s data ownership entry for the fuller treatment.

Is a data steward the same as a data custodian?

No. A steward typically holds ongoing accountability for a dataset’s quality, documentation, and lifecycle decisions; a custodian typically implements the technical controls (access permissions, backups, retention/disposal) that the steward or owner’s decisions require. In practice, especially at smaller institutions, one person or unit sometimes holds both functions, but the distinction still matters for who’s accountable when something goes wrong.

Does research data governance policy apply to data shared with external collaborators?

Yes, and this is usually where governance policy is tested most directly: cross-institutional or industry-sponsored sharing is typically gated through a formal data sharing agreement or data use agreement that encodes the access, security, and downstream-use conditions the governance policy requires.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →