Examples
Worked examples
- Is an instance
A GWAS cohort study funded by NIH deposits participant-level genotype calls and linked phenotype data to dbGaP under controlled access, while its summary association statistics and study documentation are released under open access.
- Is an instance
A secondary-analysis researcher submits a Data Access Request through dbGaP to combine two existing dbGaP-hosted cardiovascular-disease cohorts; the institution's Signing Official co-signs and the relevant Data Access Committee approves the request for a one-year term.
Counter-examples
Looks similar, but isn't
- Not an instance
A study that deposits only aggregate allele-frequency tables and de-identified summary statistics, with no participant-level genotype-phenotype linkage, does not require Data Access Committee review -- there is no controlled-access tier for data that was never individual-level to begin with.
- Not an instance
A purely clinical, non-genomic dataset (e.g. a survey-based phenotype study with no genotyping) does not belong in dbGaP regardless of funder or sensitivity; it belongs in a general-purpose or discipline-specific data repository, or a sensitive-data repository if it carries other confidentiality constraints.
Editorial commentary
dbGaP (the Database of Genotypes and Phenotypes) is the NIH/NCBI-operated repository for genotype-phenotype study data — genome-wide association study (GWAS) results, sequencing and other omics data, and the phenotypic and clinical data linked to it — submitted by NIH-funded and other studies since 2007. What distinguishes dbGaP from a general genomic-sequence archive is its two-tier access model: it deliberately separates data that is safe to release without restriction from individual-level human data that could re-identify a study participant, and gates the latter behind a formal review process rather than a login wall.
What makes something a dbGaP submission
Operationally, a study’s data lands in dbGaP — rather than an open, unrestricted repository like GenBank or a generalist repository — when it meets both of these conditions:
- It links genotype to phenotype at the individual level. A study that deposits only aggregate summary statistics (allele frequencies, association p-values with no participant-level linkage) can often stay in the open-access tier; a study that retains individual-level genotype records tied to individual-level phenotype, clinical, or demographic data cannot, because that linkage is inherently re-identifiable even after direct identifiers are stripped.
- It falls under a data-sharing expectation that names dbGaP (or an equivalent NIH-designated genomic repository) as the deposit target — most commonly the NIH Genomic Data Sharing (GDS) Policy, effective January 25, 2015, which requires NIH-funded investigators generating large-scale human or non-human genomic data (GWAS, SNP arrays, whole-genome/exome sequencing, transcriptomic, epigenomic, and related data) to submit it to an NIH-designated data repository as a term of the award.
Two access tiers: open vs. controlled
Every dbGaP study record is split across two tiers, and the split is the whole point of the repository:
- Open (unrestricted) access — anyone can browse and download without applying: the study’s public documentation (protocol, consent-form language, data dictionaries, questionnaires), summary-level variable descriptions, and non-sensitive aggregate analyses.
- Controlled access — de-identified individual-level genotype and phenotype data, pedigrees, and participant-level association results. This tier requires a successful Data Access Request (DAR) approved by the study’s Data Access Committee (DAC) before a single record can be downloaded.
Individual-level data stays de-identified in both tiers, but de-identification alone is not treated as sufficient protection for genomic data — because a genotype is inherently linkable back to a specific person given enough auxiliary information, dbGaP relies on DAC review, not just data masking, as the actual access control.
The Data Access Request (DAR) process
Getting to controlled-access data is a research-administration workflow with real institutional steps, not a self-service download:
- PI authentication via eRA Commons. The requesting Principal Investigator (and any co-investigators who need direct data access) must have an eRA Commons account and log into the dbGaP Authorized Access System with it at least once before a request can be built.
- Project request submission. The PI selects the study/dataset, states the proposed research use, and agrees to the Data Use Certification (DUC) Agreement and the Genomic Data User Code of Conduct — the terms governing confidentiality, permitted use, publication acknowledgment of the original data-generating investigators, and a prohibition on re-identification attempts.
- Institutional co-signature. The request must be co-signed by the institution’s eRA Commons Signing Official (typically in the sponsored-programs or grants office) before it reaches NIH — an institutional-certification step research administrators specifically own.
- Data Access Committee review. The relevant NIH DAC — there are many, organized by the institute or consortium that oversees a given dataset, not one central committee — reviews the request against the study’s data use limitations, the boundaries set by the original participants’ informed consent (for example, disease-specific research only, or no commercial use). Approval makes the PI an “Approved User.”
- Annual renewal or close-out. Approved-user access runs on a one-year data access period. Before it expires, the PI must submit either a Project Renewal (with a Progress Update summarizing how the data were used and any resulting publications) or a Close-out; the account is suspended if neither is submitted within 42 days of the expiration date.
dbGaP’s role in NIH genomic data sharing compliance
For research administrators, dbGaP is the operational mechanism through which two separate NIH policy obligations actually get discharged. Investigators generating large-scale genomic data under the GDS Policy satisfy their submission obligation by depositing to dbGaP (or another NIH-designated genomic repository); investigators who later want to reuse someone else’s dbGaP-deposited data for a new project satisfy their access obligation through the DAR/DAC process above, which is also the mechanism that enforces the original participants’ consent boundaries downstream. The GDS Policy sits alongside, not inside, the NIH 2023 Data Management and Sharing (DMS) Policy: a genomics-generating NIH award is typically subject to both — the DMS Policy’s general data-management-and-sharing-plan requirement, and the GDS Policy’s more specific genomic-data submission and institutional-certification requirements. A study’s Data Management and Sharing Plan for a genomics-generating award should name dbGaP (or the applicable NIH-designated repository) as the deposit target and describe how controlled-access review will be handled, not just assert that data will be “shared.”
Worked examples
- A GWAS cohort study funded by NIH deposits participant-level genotype calls and linked phenotype data (disease status, demographics) to dbGaP under controlled access, while its summary association statistics and study documentation are released under open access — satisfying the GDS Policy’s submission requirement while protecting individual participants.
- A secondary-analysis researcher at a different institution wants to combine two existing dbGaP-hosted cardiovascular-disease cohorts. Their PI submits a DAR through dbGaP describing the proposed secondary use, the institution’s Signing Official co-signs, and the relevant DAC approves the request for a one-year term, consistent with the original studies’ consent language restricting use to cardiovascular research.
Counter-example
A study that deposits only aggregate allele-frequency tables and de-identified summary statistics with no participant-level genotype-phenotype linkage — and makes no individual-level data available at all — does not require DAC review for that deposited content, because there is no controlled-access tier for data that was never individual-level to begin with. Similarly, a purely clinical dataset with no genomic component (for example, a survey-based phenotype study with no genotyping) does not belong in dbGaP regardless of funder or sensitivity; it would instead go to a general-purpose or discipline-specific data repository, or a sensitive-data repository if it carries other confidentiality constraints.
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="dbGaP (Database of Genotypes and Phenotypes)"
vocab-term-identifier="https://casrai.org/dictionary/term/dbgap-database-of-genotypes-and-phenotypes" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/dbgap-database-of-genotypes-and-phenotypes",
"name": "dbGaP (Database of Genotypes and Phenotypes)",
"identifier": "https://casrai.org/dictionary/term/dbgap-database-of-genotypes-and-phenotypes",
"description": "The NIH/NCBI repository for genotype-phenotype study data (GWAS results, sequencing/omics data, linked phenotype data), split into an open-access tier (study documentation, summary statistics) and a controlled-access tier (de-identified individual-level genotype and phenotype records) gated by Data Access Committee review of a Data Access Request -- distinct from a general sequence archive because the individual-level linkage, not just the data type, is what triggers controlled access.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
"url": "https://casrai.org/dictionary/term/dbgap-database-of-genotypes-and-phenotypes",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"dateModified": "2026-07-17T05:39:29",
"inLanguage": "en"
}






