Direct comparison
De-Identified vs. Coded vs. Anonymized Data
De-identified, coded, anonymized, pseudonymized: four terms, three legal regimes (HIPAA, Common Rule, GDPR). Compare definitions and IRB impact.
Written and maintained by CASRAI Editorial Board
Last updated
Ask CASRAI · free to try
Ask about De-Identified vs. Coded vs. Anonymized Data
Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.
An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.
Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
Works on this site and inside Claude, Cursor and the AI tools you already use.
Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.
How do De-identified (HIPAA), Coded (Common Rule), Anonymized (GDPR), Pseudonymized (GDPR) compare side by side?
The table below compares De-identified (HIPAA), Coded (Common Rule), Anonymized (GDPR), Pseudonymized (GDPR) across 7 procurement-relevant dimensions, from governing framework through where used in research.
Side-by-side comparison
| Dimension | De-identified (HIPAA) | Coded (Common Rule) | Anonymized (GDPR) | Pseudonymized (GDPR) |
|---|---|---|---|---|
| Governing framework | HIPAA Privacy Rule, 45 CFR 164.514 | Common Rule, 45 CFR 46.102, via OHRP guidance | EU GDPR, Recital 26 | EU GDPR, Article 4(5) |
| Reversible by anyone? | Covered entity may retain a secure key (164.514(c)); recipient cannot reverse it | Yes, by whoever holds the key | No, by definition | Yes, by whoever controls the key |
| Still regulated as identifiable data? | No - outside HIPAA PHI scope entirely | Depends entirely on key-holder access - the determining question | No - outside GDPR scope entirely | Yes - remains personal data, GDPR applies in full |
| Triggers IRB review? | De-identification itself is not the IRB test, but de-identified data is generally not identifiable private information | Only if investigators can readily ascertain identity via the key - the OHRP two-part test | Not a Common Rule concept | Not a Common Rule concept |
| Typical technique | Remove all 18 Safe Harbor categories, or Expert Determination risk certification | Replace identifiers with a study code; key stored separately | Suppression, generalization, k-anonymity, differential privacy, synthetic data | Coded/tokenized identifiers, encryption with a separately-stored key |
| Safe for open repository deposit? | Generally yes | Depends on key-holder terms, not the code alone | Generally yes, once genuinely anonymous | Generally no while the key remains linkable - needs a data sharing agreement |
| Where used in research | Health data leaving a covered entity for research or public release | Biospecimen repositories, longitudinal cohorts, secondary-use records | Public deposit, long-term archiving, once re-contact isn't needed | Active data collection and analysis, when re-contact/correction may be needed |
Common questions
Common questions about De-identified (HIPAA) vs Coded (Common Rule) vs Anonymized (GDPR) vs Pseudonymized (GDPR)
Is de-identified data the same as anonymized data?
+
Not necessarily. HIPAA de-identification allows the covered entity to retain a secure re-identification key (45 CFR 164.514(c)). GDPR anonymisation requires that no one, including the original data holder, can reverse the process by any reasonably likely means. A dataset can be HIPAA de-identified while still being GDPR pseudonymized if a workable key exists anywhere.
Does coding data avoid the need for IRB review?
+
Only if OHRP's conditions are met: the information or specimens already existed before the study, and investigators genuinely cannot obtain the re-identification key under a binding agreement, institutional policy, or law. This should be a documented IRB determination, not self-certified by the research team.
Can pseudonymized data be shared in an open repository?
+
Generally no while the key remains linkable, because it is still personal data under GDPR. Many studies pseudonymize during active data collection, then produce a separately anonymized derivative for public deposit once re-contact with participants is no longer needed.
Does removing a key make coded data anonymized?
+
It can move a dataset toward anonymization, but only if no other reasonably available means of re-identification remain anywhere. This should be assessed and documented, and destroying a key does not retroactively change whether IRB review was required while the key existed.
Who decides whether a dataset counts as de-identified, coded, anonymized, or pseudonymized?
+
For HIPAA de-identification: the covered entity, or an Expert Determination statistician for that method specifically. For Common Rule coded-data status: the IRB, based on the key-holder facts presented to it. For GDPR anonymisation/pseudonymisation: ultimately a legal/data-protection determination, often made with input from a Data Protection Officer, since “identifiability” under GDPR is assessed contextually, not by a fixed checklist.








