Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack BProposedv2026.1

gnomAD (Genome Aggregation Database)

gnomAD is a population-scale reference database of aggregated human exome and genome sequencing data, developed by an international coalition of investigators at the Broad Institute of MIT and Harvard, that reports allele frequencies and quality metrics for genetic variants across major ancestry groups so researchers can distinguish common polymorphisms from likely disease-causing rare variants. It is a frequency reference, not a clinical diagnostic tool or a repository of raw individual-level sequencing reads -- gnomAD releases jointly-called, harmonized summary statistics, not per-participant genomes.

ByCASRAI Editorial Board
· Last updated 1 Sept 2026
Share this

Ask about gnomAD (Genome Aggregation Database)

Answers are drawn from this dictionary entry and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Examples

Worked examples

  • Is an instance

    A clinical geneticist evaluating a novel missense variant found in a patient checks its allele frequency in gnomAD: a frequency of 30% in the general population is strong evidence against pathogenicity for a rare Mendelian disease, while absence from more than 800,000 sequenced individuals supports -- without proving -- a rare or de novo classification.

  • Is an instance

    A population geneticist studying selection against loss-of-function mutations uses gnomAD's constraint metrics (such as the LOEUF score) to identify genes depleted for predicted loss-of-function variants in the general population, flagging them as likely essential or disease-associated.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A single research lab's internal cohort of 200 exome-sequenced patients is not part of gnomAD merely because it was sequenced on a compatible platform -- gnomAD specifically means the harmonized, jointly-called aggregation the gnomAD consortium itself produces and periodically releases through its own pipeline, not any private exome or genome dataset a lab happens to hold.

Editorial commentary

The Genome Aggregation Database (gnomAD) is a population-scale reference resource of aggregated human exome and genome sequencing data, developed and maintained by an international coalition of investigators centered at the Broad Institute of MIT and Harvard. As of the current v4.1 release, gnomAD aggregates 734,947 exomes and 76,215 genomes — over 800,000 sequenced individuals in total, drawn in large part from the UK Biobank alongside many other studies — making it the largest publicly available catalogue of human genetic variation currently in use.

What gnomAD actually provides

gnomAD does not distribute raw, per-participant sequencing reads. What it releases is harmonized, jointly-called summary data: for every variant observed across its aggregated samples, gnomAD reports the allele frequency, the number of individuals carrying it, quality metrics, and — broken out by major ancestry/genetic-ancestry group — how common that variant is in each population. This population-frequency lens is what makes gnomAD indispensable for clinical variant interpretation: under the ACMG/AMP variant-classification framework, a variant’s frequency in a large, presumably healthy reference population is one of the standard lines of evidence weighed against pathogenicity.

How gnomAD differs from dbSNP and ClinVar

gnomAD, dbSNP, and ClinVar are frequently confused because all three catalogue human genetic variants, but they answer different questions. dbSNP assigns a persistent identifier to a variant regardless of frequency or clinical meaning — it is a catalogue of “does this variant exist and where.” gnomAD answers “how common is this variant, and in which populations” — it is a frequency reference built from a large, curated set of sequenced individuals. ClinVar answers “what clinical significance has been asserted for this variant” — it aggregates interpretations submitted by clinical laboratories and expert panels, not raw frequency data. A single variant routinely has entries in all three, each contributing a different piece of the interpretation picture.

Access and data use

gnomAD data is browsable and downloadable free of charge through the gnomAD browser, with bulk data available via Google Cloud, Azure, and AWS public dataset programs. Because gnomAD samples are drawn from many underlying studies whose participants consented to broad research and frequency-reporting uses (not necessarily unrestricted redistribution of individual-level data), gnomAD releases only aggregate/summary statistics publicly — individual-level genotypes are not available even to gnomAD’s own consortium members outside the original contributing studies.

Examples

  • A clinical geneticist evaluating a novel missense variant found in a patient checks its allele frequency in gnomAD: a frequency of 30% in the general population is strong evidence against pathogenicity for a rare Mendelian disease, while absence from more than 800,000 sequenced individuals supports — without proving — a rare or de novo classification.
  • A population geneticist studying selection against loss-of-function mutations uses gnomAD’s constraint metrics (such as the LOEUF score) to identify genes depleted for predicted loss-of-function variants in the general population, flagging them as likely essential or disease-associated.

Counter-example

A single research lab’s internal cohort of 200 exome-sequenced patients is not part of gnomAD merely because it was sequenced on a compatible platform — gnomAD specifically means the harmonized, jointly-called aggregation the gnomAD consortium itself produces and periodically releases through its own pipeline, not any private exome or genome dataset a lab happens to hold.

Related infrastructure

Researchers working with population-scale variant data should also be aware of the dbGaP controlled-access mechanism many of gnomAD’s contributing studies rely on for their underlying individual-level data, the Genomic Data Commons (GDC) for harmonized cancer genomics, and the earlier 1000 Genomes Project, whose public reference panel gnomAD’s scale and design build directly on.

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="gnomAD (Genome Aggregation Database)"
      vocab-term-identifier="https://casrai.org/dictionary/term/genome-aggregation-database-gnomad" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/genome-aggregation-database-gnomad",
  "name": "gnomAD (Genome Aggregation Database)",
  "identifier": "https://casrai.org/dictionary/term/genome-aggregation-database-gnomad",
  "description": "gnomAD is a population-scale reference database of aggregated human exome and genome sequencing data, developed by an international coalition of investigators at the Broad Institute of MIT and Harvard, that reports allele frequencies and quality metrics for genetic variants across major ancestry groups so researchers can distinguish common polymorphisms from likely disease-causing rare variants. It is a frequency reference, not a clinical diagnostic tool or a repository of raw individual-level sequencing reads -- gnomAD releases jointly-called, harmonized summary statistics, not per-participant genomes.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
  "url": "https://casrai.org/dictionary/term/genome-aggregation-database-gnomad",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-09-01T07:46:14",
  "dateModified": "2026-09-01T07:46:14",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →