Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Genome Projects and Research Data Management

How the Human Genome Project, 1000 Genomes, and GenBank shaped modern research data management, from the Bermuda Principles to today’s open/controlled-access repository model.

The large public genome-sequencing initiatives launched from the 1990s onward — the Human Genome Project, the 1000 Genomes Project, ENCODE, and the databases built to hold their output — are not just landmarks in molecular biology. They are the origin point for much of today’s research data management practice: the first large-scale, funder-enforced policies requiring rapid, structured, pre-publication data deposit into public repositories. A research administrator or data steward working on any funder data-sharing policy today is, in effect, working downstream of decisions genome-project consortia made about data release in 1996 and 2003.

This guide covers genome projects specifically through that lens: what data-sharing obligations these projects imposed on participating labs, which repositories they built or relied on, and how those norms carried forward into the two-tier open/controlled-access repository model that still governs human genomic data today.

What Makes a “Genome Project” Relevant to Research Data Management

For RDM purposes, a genome project is a large, typically multi-institutional, publicly or philanthropically funded initiative to sequence, catalog, or characterize the genome (or genomes) of an organism or population, with an explicit policy governing how the resulting sequence and associated data are released, deposited, and reused. Three features distinguish these projects from an individual investigator’s sequencing study:

  • Scale — output measured in gigabases or terabases, generated across multiple sequencing centers rather than a single lab.
  • A negotiated, written data-release policy — agreed on by the consortium (and often required by the funder) before sequencing began, not decided project-by-project after the fact.
  • Designation as a “community resource” — the data is produced specifically to be used by researchers outside the generating consortium, which creates a direct tension between rapid open access and the data producers’ own right to publish first analyses of their own work.

That tension — and how successive projects resolved it — is the throughline of this guide.

The Human Genome Project and the Bermuda Principles (1996)

The Human Genome Project (HGP) ran from 1990 to 2003 as an international public-sector effort, led in the United States by the NIH and the Department of Energy and internationally by centers including the Wellcome Trust Sanger Institute in the UK, alongside partners in Japan, France, Germany, and China. In February 1996, at a strategy meeting in Bermuda, leaders of the major sequencing centers adopted what became known as the Bermuda Principles: an agreement that all human genomic sequence assemblies of 1–2 kb or greater would be released to the public domain, deposited in a public database (GenBank), within 24 hours of assembly.

The NHGRI formally adopted this as a grantee data-release policy in 1997. The rationale was explicitly about keeping the emerging reference genome an unrestricted, unpatentable public resource rather than allowing any single lab or company to control access to human sequence data. Sequencing centers retained the right to publish their own first analyses of the data they generated, but the raw sequence itself had to be public almost immediately — a strikingly aggressive rapid-release requirement by the standards of research data policy at the time, and one that predates nearly every funder data-sharing mandate now in force by more than a decade.

The Fort Lauderdale Principles (2003): Rapid Release Beyond the Human Genome

As genomics expanded past the HGP itself, NHGRI convened a follow-up meeting in Fort Lauderdale in 2003 to extend the Bermuda approach to a broader category NHGRI began calling “community resource projects” — large datasets generated specifically for community-wide reuse, not just human reference sequence. The resulting Fort Lauderdale Principles kept the core commitment to rapid, pre-publication data release, while formalizing the balance that had operated informally under Bermuda: data producers could lay out their own analysis plans in advance (“marker papers”) and retained a reasonable first right to publish global analyses of the dataset, while the underlying data itself remained openly available to any researcher, immediately, for their own use.

This producer-rights-versus-open-access balance is the direct conceptual ancestor of the pre-publication data-use agreements and embargo periods that funders and consortia still write into data-sharing policies today.

The 1000 Genomes Project: Open Population-Scale Data Under Fort Lauderdale Norms

The 1000 Genomes Project (2008–2015) sequenced genomes from more than 2,500 individuals across multiple populations worldwide to catalog human genetic variation at population scale. It operated under a data-use statement explicitly following the Fort Lauderdale (and the related Toronto) principles: participant consent forms specified that sequence and genotype data would be released openly on the internet, prior to publication, while the consortium reserved first right to publish its own global analyses.

Operationally, the project was one of the first to generate multi-terabase raw sequencing datasets, which pushed the field toward dedicated sequence read archive infrastructure. Its data coordination center was run jointly by the European Bioinformatics Institute (EBI) and the U.S. National Center for Biotechnology Information (NCBI), submitting to what became the Sequence Read Archive (SRA) at both the European Nucleotide Archive (ENA) and NCBI — itself an early, working example of the mirrored, federated repository model described below.

Other Community Resource Projects: ENCODE and Beyond

NHGRI’s ENCODE (Encyclopedia of DNA Elements) project, launched in 2003 to catalog functional elements in the human genome, adopted a comparable rapid-release data policy under the same community-resource-project framework established at Fort Lauderdale. Subsequent large consortia — from international cancer genomics initiatives to model-organism sequencing projects — have generally followed the same template: negotiate a written data-release policy before generation begins, deposit into an established public or controlled-access repository, and grant the generating consortium a defined, time-limited first-publication window rather than indefinite embargo.

Where the Data Lives: GenBank, ENA, DDBJ, and the INSDC Model

Open genomic sequence data from these projects is not held in one place. It is submitted to any of three synchronized, mirrored databases that together form the International Nucleotide Sequence Database Collaboration (INSDC): GenBank, maintained by NCBI in the United States; the European Nucleotide Archive (ENA), maintained by EMBL-EBI; and the DNA Data Bank of Japan (DDBJ). A sequence submitted to any one of the three is exchanged and mirrored across all three on a daily basis, using shared accession-numbering and format conventions, so a researcher anywhere can retrieve the same record regardless of which database they query.

This is a working, decades-old precedent for the kind of federated, discipline-specific domain repository model that RDM guidance now recommends generally for structured research data: rather than one global open-data platform, disciplines with high-volume, highly structured output build dedicated, interoperable infrastructure suited to that data type. See CASRAI’s guide on how to choose an open data repository for how this discipline-specific-versus-generalist decision applies outside genomics, and the discipline-specific repository and trusted digital repository dictionary entries for the underlying concepts.

Controlled-Access Data: dbGaP, EGA, and the NIH Genomic Data Sharing Policy

Not all genomic data generated by these projects is appropriate for GenBank-style open deposit. Individual-level genotype-phenotype data — the record linking a specific participant’s genetic variants to their health data — carries re-identification risk even after de-identification, because a person’s genome is itself a unique identifier. This is why NIH built a separate, controlled-access repository, dbGaP (the Database of Genotypes and Phenotypes), established in 2007, alongside its open-access infrastructure. The European equivalent is the European Genome-phenome Archive (EGA), jointly run by EMBL-EBI and Spain’s Centre for Genomic Regulation.

NIH formalized this two-tier open/controlled-access model as explicit funder policy with the Genomic Data Sharing (GDS) Policy, issued in the NIH Guide in August 2014. The GDS Policy requires NIH-funded investigators generating large-scale human or non-human genomic data to submit it to an NIH-designated repository — dbGaP being the primary one for individual-level human genomic and phenotypic data — and requires a data-access request (DAR) process, reviewed by a Data Access Committee, for any researcher seeking to use controlled-access data, governed by a Data Use Certification agreement that specifies permitted uses and re-identification prohibitions. This is the direct funder-policy descendant of the Bermuda/Fort Lauderdale rapid-release tradition, adapted for data that cannot be released as openly as raw reference sequence: the commitment to broad reuse is preserved, but gated through access review rather than made unconditionally public. See CASRAI’s dictionary entries on sensitive-data repository and sensitive-data handling in a DMP for how this pattern generalizes to other sensitive human-subjects data.

What Genome Projects Established for Research Data Management Today

Several norms that now read as standard RDM practice were pioneered, tested, or forced into existence by genome-project data policy specifically:

  • Rapid pre-publication release as an enforceable funder condition — the Bermuda Principles predate essentially every general-purpose funder data-sharing mandate, including current NIH and NSF data management plan requirements, by well over a decade.
  • The producer-rights/open-access balance now written into most consortium and funder data-use agreements traces directly to Fort Lauderdale’s “marker paper” and first-publication-window compromise.
  • Federated, discipline-specific repository infrastructure (INSDC’s GenBank/ENA/DDBJ triad) as a proof of concept that a global, mirrored, standards-based system can outperform a single centralized platform — a model genomics reached decades before most other disciplines needed it.
  • The open/controlled-access split (GenBank versus dbGaP/EGA) as the working template funders now apply broadly whenever a dataset combines shareable derived results with re-identifiable individual-level data.
  • Structured accession identifiers as a functioning, decades-old persistent-identifier system for datasets, well before DOIs for data became standard practice — relevant background for anyone documenting data provenance that traces back through a genomic accession number.

For a research administrator writing or reviewing a data management plan involving genomic sequencing, the practical takeaway is that funder expectations in this space are unusually mature and specific compared to most other data types: know in advance whether the data is appropriate for open (GenBank/ENA/DDBJ-type) deposit or requires a controlled-access repository like dbGaP or EGA, and budget review time for the corresponding data access process. See CASRAI’s broader Research Data Management hub and the guide on making a dataset FAIR for how these genomics-specific norms connect to general-purpose FAIR and data-sharing practice.

Frequently Asked Questions

What was the Human Genome Project’s data-sharing policy?

Under the Bermuda Principles (1996), sequencing centers agreed to deposit human genomic sequence assemblies of 1–2 kb or larger into GenBank within 24 hours of assembly, keeping the reference genome an unrestricted public resource rather than allowing exclusive control or patenting of raw sequence data.

Is 1000 Genomes Project data publicly available?

Yes. The 1000 Genomes Project released its sequence and genotype data openly, prior to publication, under a data-use statement following the Fort Lauderdale and Toronto principles, with the consortium retaining first right to publish its own global analyses of the combined dataset.

What is the difference between GenBank and dbGaP?

GenBank is an open-access repository for nucleotide sequence data with no meaningful re-identification risk on its own. dbGaP is a controlled-access repository for individual-level genotype-phenotype data, where a specific person’s genetic variants are linked to their health information, requiring a reviewed data-access request and a Data Use Certification agreement before a researcher can obtain the data.

What are the Bermuda Principles?

The Bermuda Principles are the 1996 agreement among Human Genome Project sequencing centers to release human sequence assemblies of 1–2 kb or greater to public databases within 24 hours of generation, formally adopted by NHGRI as a grantee policy in 1997.

What is INSDC?

The International Nucleotide Sequence Database Collaboration is the arrangement among GenBank (NCBI, US), the European Nucleotide Archive (EMBL-EBI), and the DNA Data Bank of Japan (DDBJ) to synchronize and mirror submitted sequence data daily, so a record submitted to any one database is retrievable from all three.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →