Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

EGA (European Genome-Phenome Archive): The EU’s Controlled-Access Counterpart to dbGaP

The European Genome-phenome Archive (EGA) is the EU’s controlled-access repository for human genomic, phenotypic, and clinical research data, jointly run by EMBL-EBI and CRG. Here is how its Data Access Committee model, GDPR compliance, and Federated EGA architecture work, and how it compares to dbGaP.

The European Genome-phenome Archive (EGA) is the European Union’s controlled-access repository for human genomic, phenotypic, and clinical research data. Where dbGaP fills this role for NIH-funded and other US research, the EGA is the working default for EU-based projects that generate individual-level human genomic data too sensitive for open deposit. This guide covers what the EGA is, who governs it, what data it hosts, how its Data Access Committee model works, how it addresses GDPR, and how it compares to dbGaP.

What is the EGA?

The EGA describes itself as a global network for the permanent archiving and sharing of personally identifiable genetic, phenotypic, and clinical data generated for biomedical research, including research-focused healthcare systems. It launched in 2008 at the European Bioinformatics Institute (EMBL-EBI) to address a specific archiving gap: genome-wide association study (GWAS) data and other individual-level genomic datasets that could re-identify a study participant, and therefore could not be deposited in an open, unrestricted sequence archive the way aggregate or non-identifiable data can.

The EGA is a recognised driver project of the Global Alliance for Genomics and Health (GA4GH), the international standards body that develops technical and policy frameworks for responsible genomic and health data sharing.

Who runs the EGA: EMBL-EBI and CRG

The EGA is jointly operated by two organisations:

  • EMBL-EBI (the European Bioinformatics Institute), part of the European Molecular Biology Laboratory, based in Hinxton, near Cambridge, UK.
  • CRG (the Centre for Genomic Regulation), a research institute based in Barcelona, Spain.

The two institutions formalised their joint governance of the archive in stages: a memorandum of understanding signed in 2013, followed by a formal joint agreement in 2016 that established EGA as a shared project rather than an EMBL-EBI-only service. This joint UK/Spain operating structure is a genuine point of contrast with dbGaP, which is run by a single US federal agency (NIH, via NCBI) rather than a cross-border partnership.

What data the EGA hosts

The EGA hosts genomic, phenotypic, and clinical datasets generated by biomedical research, wherever that data is sensitive enough that participants could be re-identified from it. In practice this covers the same broad categories dbGaP handles for US studies: genome-wide association study (GWAS) genotype calls, whole-genome and whole-exome sequencing data, and other omics data, together with the phenotypic, clinical, and demographic records linked to specific participants. Study- and dataset-level metadata (titles, descriptions, and accession numbers) is discoverable publicly through the EGA’s own search interface even though the underlying data files themselves sit behind controlled access — the same open-metadata/controlled-data split used by dbGaP and described in CASRAI’s guide to genome projects and research data management.

Controlled access and Data Access Committees (DACs)

Access to each individual study or dataset in the EGA is controlled by that study’s own Data Access Committee (DAC) — a body of one or more named individuals, typically drawn from the organisation that originally collected the samples and obtained participant consent, responsible for deciding whether a specific data access request can be granted. A DAC evaluates each request against the terms of the original participant consent and any applicable national research-ethics approval, not against a single archive-wide policy; this means access conditions can genuinely differ from one EGA study to the next, unlike a single centralised review body.

A researcher requesting controlled-access EGA data typically needs to: register for an EGA account, identify the specific study or dataset and locate its DAC contact, submit a request explaining the intended research use, and, if approved, agree to a Data Access Agreement that sets out permitted uses and security obligations before the data is released. This per-study DAC model is the EGA’s structural analogue to dbGaP’s Data Access Committee and Data Use Certification process, though the two systems are administered differently in practice (see the comparison below).

Related governance models exist outside Europe and the US too — see CASRAI’s dictionary entry on H3Africa’s Data and Biospecimen Access Committee (DBAC) for how a comparable controlled-access model was adapted for African genomic research.

GDPR and EU data protection compliance

Because the EGA is a European infrastructure handling data that identifies or can re-identify living individuals, its Data Access Committees are explicitly tasked with ensuring that data release complies with the EU General Data Protection Regulation (GDPR), alongside the terms of participant consent. This is a structural difference from dbGaP, which operates under US federal policy (principally the NIH Genomic Data Sharing Policy) rather than the GDPR. In practice this means an EU-based study depositing to the EGA, or a researcher anywhere in the world requesting EGA-held data on EU participants, needs to account for GDPR’s requirements around lawful basis, purpose limitation, and cross-border transfer of personal data as part of the access process, not just the archive’s own data use terms. CASRAI’s dictionary covers the specific lawful-basis question many publicly funded studies rely on in the entry on the GDPR Article 6(1)(e) public task legal basis.

The Federated EGA (FEGA)

Since 2022, the EGA has been transitioning from a single centralised archive toward the Federated EGA (FEGA), a network in which participating countries operate their own EGA-compliant national nodes rather than depositing everything into one central instance. The initiative launched with founding nodes in Finland, Germany, Norway, Spain, and Sweden, with additional national nodes joining since. The aim of federation is to let sensitive genomic data remain hosted within the jurisdiction (and under the data protection regime) where it originated, while still being discoverable and requestable through a harmonised, EGA-compatible submission and access process across the network. This is a meaningful architectural difference from dbGaP, which remains a single centralised US repository.

EGA vs. dbGaP: how the two controlled-access models compare

  EGA dbGaP
Operated by EMBL-EBI (UK) and CRG (Spain), jointly NIH/NCBI (United States)
Launched 2008 (joint EMBL-EBI/CRG agreement, 2016) 2007
Access review body Per-study/dataset Data Access Committee (DAC), typically the originating institution Data Access Committee reviewing a Data Access Request (DAR) under a Data Use Certification agreement
Primary data-protection framework EU GDPR plus participant consent/ethics terms US NIH Genomic Data Sharing (GDS) Policy plus participant consent/ethics terms
Architecture Transitioning to a federated network of national nodes (FEGA, from 2022) Single centralised US repository
Open-access counterpart Public study/dataset metadata; underlying files controlled Open and controlled tiers within the same study record

The similarity that matters most for research administrators is structural, not procedural: both repositories implement the same underlying principle — individual-level human genomic and phenotypic data that could re-identify a participant is not appropriate for open deposit, and needs a governed request-and-approval process instead of a login wall. Which repository a given study uses is generally determined by funder mandate and the location of the originating institution, not researcher preference; some international consortia submit to both, or to a discipline-specific repository that itself interfaces with EGA or dbGaP.

How EGA fits into a data management plan

For an EU-based (or EU-participant) study generating individual-level human genomic data, naming the EGA as the intended controlled-access repository in a data management plan is only the first step; a credible DMP also needs to identify who will constitute or liaise with the study’s DAC, confirm the GDPR lawful basis under which the original consent was obtained, and budget realistic time for the access-request review process, since DAC review timelines are set by the individual committee rather than a fixed archive-wide service level. See CASRAI’s broader Research Data Management hub and the guide on genome projects and research data management for how this fits into the wider open/controlled-access planning process, and the trusted digital repository entry for the certification concepts a controlled-access archive like the EGA is expected to meet.

Frequently asked questions

What is the EGA used for?

The EGA is used to permanently archive and share human genomic, phenotypic, and clinical research data that is too sensitive or re-identifiable for open deposit, giving researchers a governed way to discover and request access to it while protecting participant privacy.

Who can access EGA controlled-access data?

Any researcher can request access, but release is decided case-by-case by the relevant study’s Data Access Committee, which evaluates the request against the terms of participant consent, applicable research-ethics approval, and GDPR before granting access under a Data Access Agreement.

Is EGA the same as dbGaP?

No. They are separate repositories serving an equivalent role for different jurisdictions: dbGaP is operated by the US NIH/NCBI, while the EGA is jointly operated by EMBL-EBI (UK) and CRG (Spain) and is governed with respect to EU GDPR rather than US federal policy. Both implement the same open/controlled-access principle for individual-level human genomic data.

Does the EGA comply with GDPR?

The EGA’s Data Access Committees are explicitly responsible for ensuring that data release complies with the GDPR, in addition to the terms of the original participant consent, since the archive routinely handles personal data of EU residents.

What is the Federated EGA (FEGA)?

FEGA is the EGA’s ongoing transition, begun in 2022, from a single centralised archive to a federated network of national nodes, allowing participating countries to host sensitive genomic data within their own jurisdiction while remaining discoverable and accessible through a harmonised EGA-compatible process.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →