Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack BProposedv2026.1

1000 Genomes Project

The 1000 Genomes Project was a completed (2008-2015) international research consortium that sequenced the genomes of 2,504 individuals across 26 populations worldwide, specifically to catalogue common human genetic variation (variants at roughly 1% or greater population frequency), and discovered more than 88 million variants in doing so. It is not an ongoing project -- its data and reference resources are now maintained and updated by the International Genome Sample Resource (IGSR) -- and it did not attempt to catalogue rare, individually clinically actionable variants the way a targeted clinical sequencing effort would.

ByCASRAI Editorial Board
· Last updated 1 Sept 2026
Share this

Ask about 1000 Genomes Project

Answers are drawn from this dictionary entry and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Examples

Worked examples

  • Is an instance

    A population geneticist studying allele-frequency differences across ancestry groups downloads 1000 Genomes Project VCF files from IGSR to use as a public reference panel for imputing untyped variants in a genotyping-array study.

  • Is an instance

    A bioinformatics tool developer benchmarks a new variant-calling pipeline against the 1000 Genomes Project's well-characterized, multiply-validated variant set precisely because its ground truth is independently confirmed and stable, unlike a live, still-changing clinical dataset.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A rare, disease-causing variant found in only one family is not the kind of variant the 1000 Genomes Project set out to catalogue -- the project specifically targeted variants at or above roughly 1% population frequency, so a rare variant's absence from the 1000 Genomes dataset says nothing about its real-world frequency or pathogenicity.

Editorial commentary

The 1000 Genomes Project (1KGP) was an international research effort, running from 2008 to 2015, to build the most detailed catalogue of common human genetic variation available at the time. Researchers sequenced the genomes of 2,504 individuals from 26 populations across Africa, the Americas, East Asia, Europe, and South Asia, and discovered more than 88 million genetic variants — single nucleotide polymorphisms, short insertions/deletions, and structural variants — in the process. All sequence data was made freely available to the global research community as it was generated, one of the earliest large-scale demonstrations of fully open genomic-data release.

What the project deliberately targeted

The 1000 Genomes Project set out specifically to catalogue common variation — variants present at roughly 1% frequency or higher across the populations studied — rather than to find every rare or private variant in any one individual’s genome. That design choice is what makes the resulting dataset so useful as a background reference panel: because it captures the common variation that most people share, it lets researchers distinguish a genuinely rare, potentially significant variant in a new sample from an already-common, likely benign one.

Where the project’s data lives now

The 1000 Genomes Project itself concluded in 2015, but its data did not stop being used. The International Genome Sample Resource (IGSR) now maintains, updates, and re-releases the project’s data collections and reference resources — including realigning the original samples to newer human reference genome builds as those are released, so the dataset stays usable by modern pipelines years after the original sequencing was completed. The project’s catalogued SNPs and short indels were also submitted to dbSNP, and its structural variant calls to the companion Database of Genomic Structural Variation (DGVa).

Examples

  • A population geneticist studying allele-frequency differences across ancestry groups downloads 1000 Genomes Project VCF files from IGSR to use as a public reference panel for imputing untyped variants in a genotyping-array study.
  • A bioinformatics tool developer benchmarks a new variant-calling pipeline against the 1000 Genomes Project’s well-characterized, multiply-validated variant set precisely because its ground truth is independently confirmed and stable, unlike a live, still-changing clinical dataset.

Counter-example

A rare, disease-causing variant found in only one family is not the kind of variant the 1000 Genomes Project set out to catalogue — the project specifically targeted variants at or above roughly 1% population frequency, so a rare variant’s absence from the 1000 Genomes dataset says nothing about its real-world frequency or pathogenicity.

Related infrastructure

See dbSNP, where the project’s short variants were catalogued, and gnomAD, the much larger, ongoing successor-scale resource for population allele frequencies today.

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="1000 Genomes Project"
      vocab-term-identifier="https://casrai.org/dictionary/term/1000-genomes-project" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/1000-genomes-project",
  "name": "1000 Genomes Project",
  "identifier": "https://casrai.org/dictionary/term/1000-genomes-project",
  "description": "The 1000 Genomes Project was a completed (2008-2015) international research consortium that sequenced the genomes of 2,504 individuals across 26 populations worldwide, specifically to catalogue common human genetic variation (variants at roughly 1% or greater population frequency), and discovered more than 88 million variants in doing so. It is not an ongoing project -- its data and reference resources are now maintained and updated by the International Genome Sample Resource (IGSR) -- and it did not attempt to catalogue rare, individually clinically actionable variants the way a targeted clinical sequencing effort would.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
  "url": "https://casrai.org/dictionary/term/1000-genomes-project",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-09-01T07:46:16",
  "dateModified": "2026-09-01T07:46:16",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →