Direct comparison
De-Identified vs. Coded vs. Anonymized Data
De-identified, coded, anonymized, pseudonymized: four terms, three legal regimes (HIPAA, Common Rule, GDPR). Compare definitions and IRB impact.
Ask about De-Identified vs. Coded vs. Anonymized Data
Answers are drawn from this comparison and the rest of the CASRAI corpus, with a link to every source.
Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this
How do De-identified (HIPAA), Coded (Common Rule), Anonymized (GDPR), Pseudonymized (GDPR) compare side by side?
The table below compares De-identified (HIPAA), Coded (Common Rule), Anonymized (GDPR), Pseudonymized (GDPR) across 7 procurement-relevant dimensions, from governing framework through where used in research.
Side-by-side comparison
| Dimension | De-identified (HIPAA) | Coded (Common Rule) | Anonymized (GDPR) | Pseudonymized (GDPR) |
|---|---|---|---|---|
| Governing framework | HIPAA Privacy Rule, 45 CFR 164.514 | Common Rule, 45 CFR 46.102, via OHRP guidance | EU GDPR, Recital 26 | EU GDPR, Article 4(5) |
| Reversible by anyone? | Covered entity may retain a secure key (164.514(c)); recipient cannot reverse it | Yes, by whoever holds the key | No, by definition | Yes, by whoever controls the key |
| Still regulated as identifiable data? | No - outside HIPAA PHI scope entirely | Depends entirely on key-holder access - the determining question | No - outside GDPR scope entirely | Yes - remains personal data, GDPR applies in full |
| Triggers IRB review? | De-identification itself is not the IRB test, but de-identified data is generally not identifiable private information | Only if investigators can readily ascertain identity via the key - the OHRP two-part test | Not a Common Rule concept | Not a Common Rule concept |
| Typical technique | Remove all 18 Safe Harbor categories, or Expert Determination risk certification | Replace identifiers with a study code; key stored separately | Suppression, generalization, k-anonymity, differential privacy, synthetic data | Coded/tokenized identifiers, encryption with a separately-stored key |
| Safe for open repository deposit? | Generally yes | Depends on key-holder terms, not the code alone | Generally yes, once genuinely anonymous | Generally no while the key remains linkable - needs a data sharing agreement |
| Where used in research | Health data leaving a covered entity for research or public release | Biospecimen repositories, longitudinal cohorts, secondary-use records | Public deposit, long-term archiving, once re-contact isn't needed | Active data collection and analysis, when re-contact/correction may be needed |
Common questions
Common questions about De-identified (HIPAA) vs Coded (Common Rule) vs Anonymized (GDPR) vs Pseudonymized (GDPR)
Is de-identified data the same as anonymized data?
+
Not necessarily. HIPAA de-identification allows the covered entity to retain a secure re-identification key (45 CFR 164.514(c)). GDPR anonymisation requires that no one, including the original data holder, can reverse the process by any reasonably likely means. A dataset can be HIPAA de-identified while still being GDPR pseudonymized if a workable key exists anywhere.
Does coding data avoid the need for IRB review?
+
Only if OHRP's conditions are met: the information or specimens already existed before the study, and investigators genuinely cannot obtain the re-identification key under a binding agreement, institutional policy, or law. This should be a documented IRB determination, not self-certified by the research team.
Can pseudonymized data be shared in an open repository?
+
Generally no while the key remains linkable, because it is still personal data under GDPR. Many studies pseudonymize during active data collection, then produce a separately anonymized derivative for public deposit once re-contact with participants is no longer needed.
Does removing a key make coded data anonymized?
+
It can move a dataset toward anonymization, but only if no other reasonably available means of re-identification remain anywhere. This should be assessed and documented, and destroying a key does not retroactively change whether IRB review was required while the key existed.
Who decides whether a dataset counts as de-identified, coded, anonymized, or pseudonymized?
+
For HIPAA de-identification: the covered entity, or an Expert Determination statistician for that method specifically. For Common Rule coded-data status: the IRB, based on the key-holder facts presented to it. For GDPR anonymisation/pseudonymisation: ultimately a legal/data-protection determination, often made with input from a Data Protection Officer, since “identifiability” under GDPR is assessed contextually, not by a fixed checklist.








