Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Managing Participant-Level Research Data

How to de-identify, tier access to, and govern sharing of participant-level (human-subjects) research data under HIPAA and GDPR, distinct from informed consent and IRB review.

“Participant data” — the records generated by or about the people enrolled in a study — is research data with an extra constraint layered on top of the usual data-management questions: it is also personal data about identifiable individuals. That changes what can be documented in a data management plan (DMP), what “sharing” is allowed to mean, and which repositories are even eligible destinations. This guide covers the research-data-management side of that problem: how to de-identify participant-level data appropriately, what access model to plan for, and what needs to be in a data use agreement — not the separate questions of how to obtain informed consent or how to protect a vulnerable population during recruitment, which are ethics-review and human-subjects-protection topics covered elsewhere (see IRB/REC Approval Process and the OHRP dictionary entry).

What makes participant-level data different in a DMP

A standard data management plan answers what data will be produced, how it will be organized and documented, where it will be stored, and how (and whether) it will be shared. Participant-level data forces a more conditional version of that last answer. A funder’s default expectation — deposit in a repository, make it as open as possible — has to be qualified by what the consent participants actually gave, what identifiers the dataset still carries, and what legal framework (HIPAA, GDPR, or an institutional equivalent) governs it. NIH’s Data Management and Sharing Policy, effective since January 25, 2023, is explicit on this point: it requires a DMS Plan for essentially all NIH-funded research producing scientific data, but it also directs investigators toward controlled-access repositories and permits justified limits on sharing where privacy constraints require it — “maximally open” is the default posture, not an unconditional mandate. Building the de-identification and access-tier decision into the DMP at proposal stage, rather than retrofitting it at the point of deposit, is what keeps a project’s actual sharing practice consistent with what was disclosed to participants and to the funder.

The de-identification spectrum: HIPAA, GDPR, and what “de-identified” doesn’t mean

“De-identified” is not one fixed state — it’s a claim that has to be operationalized against a specific standard, and the major frameworks don’t use the same one. CASRAI’s Dictionary defines de-identification operationally as removing or altering direct and indirect identifiers so that a specific individual cannot reasonably be identified, alone or in combination with other information a realistic recipient could obtain — but “reasonably” is doing real work in that sentence, and the frameworks differ on how to satisfy it:

  • HIPAA (45 CFR 164.514(a)-(b)) recognizes two concrete methods that remove data from the Privacy Rule’s scope entirely: Safe Harbor (removing eighteen specified identifier categories — names, exact dates, geographic subdivisions smaller than a state, and so on) and Expert Determination (a qualified statistician applies accepted methods to show the risk of re-identification is very small). Safe Harbor is mechanical and auditable; Expert Determination is more flexible but requires documented statistical justification.
  • GDPR draws a sharper line and never uses the term “de-identification” at all. Article 4(5) defines pseudonymisation — replacing identifiers with a key that could be used to re-link the data — as a safeguard that still counts as personal data (Recital 26), because re-identification remains possible with the key. Only genuinely anonymised data, where re-identification is not reasonably possible by anyone, falls outside GDPR’s scope entirely. Article 89(1) explicitly names pseudonymisation as an accepted safeguard for research processing — it’s a risk-reduction measure the GDPR endorses, not a way to exit the regulation. See Pseudonymisation.

The practical consequence: a dataset that satisfies HIPAA Safe Harbor is not automatically GDPR-anonymous, and a GDPR-pseudonymised dataset is not the same thing as a HIPAA-de-identified one. Whichever standard applies to your data, treat the resulting classification as a starting point for choosing an access tier (below), not as a guarantee that open deposit is now safe. Where re-identification risk can’t be brought low enough for open release — small samples, rare conditions, genomic data — the RDM answer is usually a controlled-access repository rather than heavier redaction of an open one.

Access tiers: open, registered, and controlled access

Participant-level datasets rarely fit a single open/closed binary. The tiered model used by NIH’s dbGaP repository is a useful reference structure because it makes the trade-off explicit:

  • Open access — aggregate statistics, variable dictionaries, and documentation that carry no meaningful re-identification risk are released with no request process.
  • Controlled access — de-identified individual-level records are released only to investigators who submit a Data Access Request, agree to a Data Use Certification, and are approved by a Data Access Committee for a specified research use.

A DMP that plans for participant-level sharing should state which tier each data product will fall into and why, rather than defaulting to “data available upon request” — a phrase reviewers and funders increasingly read as a placeholder rather than an actual access mechanism. If the underlying consent only covers a specific research use, controlled access with use-restricted approval is often the only model consistent with what participants agreed to, independent of the technical de-identification standard applied.

Data use agreements

Controlled-access participant data is normally released under a data use agreement (DUA) — a contract between the data provider (or repository) and the requesting institution or investigator that specifies what the recipient signs up to before receiving the data. A DUA typically restricts the data to the approved research use, prohibits attempts at re-identification, sets security and storage requirements at the recipient site, restricts redisclosure to unapproved third parties, and specifies destruction or return of the data at project end. dbGaP’s own Data Use Certification Agreement, for example, requires the requesting investigator to notify the relevant Data Access Committee of any unauthorized sharing, security breach, or inadvertent release within a short window of discovery, followed by a fuller written incident report. Whether or not you’re using dbGaP specifically, a DUA is a distinct instrument from a Certificate of Confidentiality (a legal shield against compelled disclosure, independent of any contract) and from the de-identification work itself — a dataset can be properly de-identified and still require a DUA if it’s released through a controlled-access channel, because the DUA is governing the recipient’s obligations, not the data’s technical state.

Choosing a repository for participant-level data

Not every general-purpose data repository accepts identifiable or sensitive human-subjects data, and depositing it somewhere that isn’t built for controlled access is itself a governance failure independent of how well the data was de-identified. When selecting a destination, check: whether the repository offers a controlled-access tier with a review process (not just open deposit); whether it’s named or implicitly endorsed in relevant funder guidance for your data type (dbGaP for human genotype-phenotype data under NIH’s Genomic Data Sharing Policy is the clearest example); and whether its own data use agreement terms are compatible with what your consent and institutional data-sharing agreement actually permit. Where genuinely open sharing is the goal and full participant-level records can’t be released, consider whether a synthetic version of the dataset — generated to mimic the statistical structure without corresponding to any real participant — can satisfy the sharing intent; see Synthetic Data in Research for what current funder and regulator guidance actually requires before treating synthetic data as a substitute for de-identification.

Consent language and future data sharing — the RDM interface

Informed consent itself is an ethics-review and IRB/REC matter, not an RDM one, and this guide doesn’t cover how to design or administer a consent process. But the specific wording a consent form uses about future data use and sharing directly constrains what a DMP can promise. Narrow, study-specific consent language can make broad public deposit inconsistent with what participants agreed to, even after full de-identification; broader consent for future secondary use by other researchers is what makes controlled-access sharing viable at all. Confirming, at DMP-drafting stage, what the approved consent language actually permits — rather than assuming de-identification alone clears the way to open sharing — is the point where data management planning and human-subjects protection have to be coordinated, even though they remain separate reviews with separate owners.

A practical checklist for a participant-data DMP

  • Identify which framework governs the data’s identifiability status (HIPAA, GDPR, an institutional policy, or more than one) and which de-identification method you’ll apply and document.
  • Confirm what the approved consent language permits for future sharing and secondary use, and keep the DMP consistent with it.
  • Assign each data product (raw, de-identified, aggregate/summary) to an access tier — open, registered, or controlled — rather than a single blanket sharing statement.
  • Select a repository that explicitly supports the access tier you need, not just general-purpose deposit.
  • Specify the data use agreement terms that will govern controlled-access release, including breach/incident reporting obligations.
  • Document data ownership and stewardship responsibilities separately from access control — see Data Ownership, since who controls access decisions and who legally owns the data are not always the same question.

For worked examples of how these decisions get written into an actual DMP narrative, see Data Management Plan Worked Examples.

Frequently asked questions

Is de-identified participant data still “personal data” under GDPR?

Only if it’s pseudonymised rather than fully anonymised. GDPR Recital 26 treats pseudonymised data — where a key could still be used to re-link records to individuals — as personal data still within scope, while genuinely anonymised data (no reasonable possibility of re-identification by anyone) falls outside GDPR entirely. Most practically de-identified research datasets are pseudonymised in this sense, not anonymised.

Does HIPAA de-identification automatically satisfy GDPR?

No. The two frameworks use different tests, and satisfying one doesn’t establish the other. Treat them as separate compliance questions if your data or collaborators are subject to both.

What’s the difference between a data use agreement and a Certificate of Confidentiality?

A DUA is a contract governing what an approved recipient may do with data they’ve been granted access to. A Certificate of Confidentiality is a legal protection against compelled disclosure (e.g., in legal proceedings) that applies independent of any contract. They address different risks and commonly coexist on the same study.

Can I just say “data available upon request” in my DMP?

Funders increasingly expect more than that phrase alone. State the actual access tier (open, registered, or controlled), the repository or mechanism through which a request would be processed, and the criteria a request would be evaluated against.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →