The HIPAA Privacy Rule recognizes two ways to strip a dataset of “individually identifiable health information” so that what remains falls outside the Rule’s protections: the Safe Harbor method, which works by removing a fixed list of 18 identifier categories, and the Expert Determination method, which works by having a qualified statistician certify that re-identification risk is very small. Both are defined in the same regulation, 45 CFR §164.514(b), and a covered entity or its business associate may use either one — there is no requirement to prefer Safe Harbor.
Last verified 16 August 2026 against the current text of 45 CFR §164.514 on eCFR. This page explains the regulatory de-identification standard for research and operational use; it is not legal advice, and institutional Privacy Officers or IRBs should be the final word on a specific dataset.
Safe Harbor vs. Expert Determination, at a glance
| Safe Harbor | Expert Determination | |
|---|---|---|
| Regulatory basis | 45 CFR 164.514(b)(2) | 45 CFR 164.514(b)(1) |
| How it works | Remove all 18 specified identifier categories from the dataset | A qualified expert applies statistical/scientific methods and determines re-identification risk is “very small” |
| Who signs off | The covered entity itself (no external expert required) | “A person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods” |
| Extra condition | Covered entity must have no actual knowledge remaining data could identify someone, alone or combined with other available information | None beyond the risk determination itself |
| Documentation | Not explicitly required by the rule, but recommended practice | Required — must document the methods and results that justify the “very small risk” conclusion |
| Good fit for | Structured datasets (registries, EHR extracts, claims data) where the 18 fields are known and removable | Datasets where removing all 18 categories would destroy analytic value — e.g. genomic data, free text, geospatial or longitudinal data where dates and location carry the signal |
| Re-identification risk standard | Implicit in the fixed list; not independently quantified | Explicitly quantified and justified in writing by the expert |
The 18 HIPAA Safe Harbor identifiers
Under 45 CFR §164.514(b)(2)(i), a covered entity may treat health information as de-identified under Safe Harbor only if all 18 of the following identifiers — of the individual or of the individual’s relatives, employers, or household members — are removed, and the entity has no actual knowledge that the remaining information could be used, alone or in combination, to identify the individual.
| # | Identifier category | Notes |
|---|---|---|
| 1 | Names | All name fields, including relatives, employers, household members |
| 2 | Geographic subdivisions smaller than a state | Street address, city, county, precinct, zip code and equivalent geocodes. Exception: the first 3 digits of a ZIP code may be kept if the Census-defined geographic unit sharing those 3 digits has >20,000 people; otherwise those 3 digits must be changed to 000 |
| 3 | All elements of dates (except year) directly related to an individual | Birth date, admission date, discharge date, date of death. All ages over 89, and any date elements indicating such an age, must also be removed — these may only be aggregated into a single “90 or older” category |
| 4 | Telephone numbers | |
| 5 | Fax numbers | |
| 6 | Email addresses | |
| 7 | Social Security numbers | |
| 8 | Medical record numbers | |
| 9 | Health plan beneficiary numbers | |
| 10 | Account numbers | |
| 11 | Certificate/license numbers | |
| 12 | Vehicle identifiers and serial numbers | Including license plate numbers |
| 13 | Device identifiers and serial numbers | |
| 14 | Web URLs | |
| 15 | IP addresses | |
| 16 | Biometric identifiers | Including finger and voice prints |
| 17 | Full-face photographs and comparable images | |
| 18 | Any other unique identifying number, characteristic, or code | The catch-all category — covers things like unusual occupations, rare diagnoses combined with small geography, or any other field that could function as a de facto identifier. It does not include a re-identification code the covered entity itself assigns and keeps secure under 164.514(c) |
Source: 45 CFR §164.514(b)(2)(i)(A)–(R), eCFR, current version.
The “no actual knowledge” condition people miss
Removing all 18 categories is necessary but not sufficient for Safe Harbor. 164.514(b)(2)(ii) adds a second condition: the covered entity must not have actual knowledge that the remaining information could be used, alone or combined with other available information, to identify the person. This is why Safe Harbor can still fail on a small or unusual population — a rural clinic with one patient of a rare condition can strip all 18 fields and still know, as a matter of fact, who the record describes. In that situation Safe Harbor does not apply even though the checklist was followed, and Expert Determination (or simply not sharing the record) is the appropriate route.
The Expert Determination method
45 CFR §164.514(b)(1) allows de-identification through a person “with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable” who:
- Applies those methods and determines the risk is very small that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify a subject of the information; and
- Documents the methods and results of that analysis in enough detail to justify the determination.
HHS guidance describes this as a flexible, risk-based process rather than a fixed checklist: the expert weighs who the anticipated recipients are, what other data they could realistically combine with the dataset, and how the specific variables retained (dates, granular geography, rare conditions, genomic data) affect re-identification risk in that context. There is no HHS certification or credential that makes someone an “expert” by default — the qualification is judged by actual statistical/scientific competence and a track record the covered entity can defend if questioned.
Expert Determination is the method research teams reach for when Safe Harbor’s blanket removals would destroy the dataset’s research value — for example, keeping full dates for a longitudinal analysis, exact ages over 89 for a geriatrics study, or granular geography for an environmental-exposure study, each justified and documented rather than simply excluded.
Which method to use: a quick decision check
- Use Safe Harbor when the 18 categories can be removed without destroying the analysis (most administrative, billing, and structured-EHR extracts), and you have no independent reason to believe the remainder is still identifying.
- Use Expert Determination when the research question depends on a field Safe Harbor would force you to strip — exact dates, fine-grained geography, rare-disease flags, genomic or imaging data — and you can engage a qualified statistician to assess and document risk.
- Neither method applies if you plan to disclose direct identifiers alongside clinical data for a defined research purpose with a data use agreement in place — that is a Limited Data Set, a separate, narrower HIPAA category that still counts as PHI and requires a signed data use agreement, not full de-identification.
- Check institutional policy before either. Many IRBs and Privacy Offices require documentation of which method was used and by whom, independent of the federal minimum.
De-identified data vs. a Limited Data Set
These are frequently conflated but regulatorily distinct. De-identified data (via Safe Harbor or Expert Determination) is, by definition, no longer PHI and falls outside the Privacy Rule entirely. A Limited Data Set is still PHI — it has had direct identifiers removed but may retain dates and geographic detail broader than a Limited Data Set’s own permitted fields — and can only be disclosed for research, public health, or health care operations under a signed data use agreement. See Limited Data Set vs. De-Identified Data for the full field-by-field comparison.
De-identification under HIPAA is also a distinct concept from GDPR’s anonymisation/pseudonymisation split, which uses different tests and, unlike HIPAA, has no fixed 18-field list. See Anonymization vs. Pseudonymization and Data Anonymisation in Research for how the GDPR framing differs.
Common pitfalls
- Treating Safe Harbor as purely mechanical. The “no actual knowledge” condition (164.514(b)(2)(ii)) is a real, independent requirement, not a formality.
- Forgetting the ZIP code carve-out is conditional. Only the first 3 digits qualify, and only if the combined population of that 3-digit area exceeds 20,000 per the current Census data — otherwise even those 3 digits must be zeroed out.
- Missing that ages over 89 must be aggregated. Not just birth dates — any date element indicating an age over 89 has to collapse into a single “90 or older” bucket.
- Assuming a re-identification key defeats de-identification. 164.514(c) explicitly permits the covered entity to keep a secure code for its own future re-identification, provided the code isn’t derived from the individual’s information and isn’t disclosed.
- Confusing a Limited Data Set with de-identified data because both sound like “less identifiable” — only de-identified data is outside HIPAA’s scope; a Limited Data Set is not.
Frequently asked questions
What are the 18 HIPAA identifiers?
They are the 18 categories of information listed in 45 CFR §164.514(b)(2)(i) that must all be removed from a dataset for it to qualify as de-identified under the Safe Harbor method — names, geographic subdivisions smaller than a state, most date elements, contact details, government and account numbers, device and vehicle identifiers, web/IP addresses, biometric identifiers, full-face photos, and a catch-all for any other unique identifying number, characteristic, or code. The full list with regulatory notes is in the table above.
Is de-identified data still considered PHI?
No. Health information that meets either the Safe Harbor or Expert Determination standard is, by the text of 45 CFR §164.514(a), no longer “individually identifiable health information,” so it falls outside the HIPAA Privacy Rule’s protections. A Limited Data Set, by contrast, is still PHI.
Who is qualified to perform an Expert Determination?
The regulation does not name a credential; it requires “a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable.” In practice this is typically a biostatistician or data-privacy specialist whose determination and methodology are documented and defensible if challenged.
Can a dataset include ZIP codes and still be Safe Harbor de-identified?
Only the first three digits, and only where the Census-defined population sharing those three digits exceeds 20,000 people; otherwise those three digits must be changed to 000. Full five-digit ZIP codes are never permitted under Safe Harbor.
Does removing the 18 identifiers guarantee a dataset can’t be re-identified?
Not by itself. Safe Harbor also requires that the covered entity have no actual knowledge the remaining data could still identify someone, and even government agencies have published re-identification research on datasets stripped of direct identifiers when unusual combinations of remaining variables (e.g. rare condition plus small area) still narrowed down to one person. This is exactly the scenario Expert Determination’s quantified risk analysis is designed to catch.
Does the 18-identifier list apply outside HIPAA, e.g. to non-covered-entity research data?
The Safe Harbor list is a HIPAA Privacy Rule standard and technically governs covered entities and their business associates. It is nonetheless widely adopted as a de facto best-practice checklist by non-HIPAA-covered researchers and repositories because it is specific, enumerated, and well understood — see De-identification for how the concept is operationalized across different frameworks including GDPR.
Related pages
- De-identification (dictionary)
- Limited Data Set (HIPAA)
- PHI Exemptions From the HIPAA Privacy Rule
- HIPAA and Retrospective Research
- Limited Data Set vs. De-Identified Data
- HIPAA Authorization vs. Informed Consent
- Secondary Use of Identifiable Data and Biospecimens
- Retrospective Chart Review: IRB Requirements, Exemptions, and Consent Waivers







