Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack BProposedv2026.1

ENCODE (Encyclopedia of DNA Elements)

An NHGRI-funded international research consortium, launched in September 2003, that systematically catalogs the functional and regulatory elements of the human genome (and, since ENCODE3, the mouse genome) using standardized assays -- ChIP-seq for transcription-factor binding and histone marks, RNA-seq for transcription, DNase-seq/ATAC-seq for chromatin accessibility, and Hi-C for 3D genome contacts -- run across a wide panel of cell types and tissues. Data is released through the ENCODE Portal under standardized accessions (ENCSR for experiments, ENCFF for files), validated and harmonized by a central Data Coordination Center, and no dataset is 'ENCODE data' unless it was actually submitted to and processed through that pipeline.

ByCASRAI Editorial Board
· Last updated 1 Sept 2026
Share this

Ask about ENCODE (Encyclopedia of DNA Elements)

Answers are drawn from this dictionary entry and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Examples

Worked examples

  • Is an instance

    A regulatory-genomics researcher downloads a bigWig ChIP-seq signal track for a specific transcription factor in a liver cell line from the ENCODE Portal, citing the file's ENCFF accession in the manuscript's data-availability statement.

  • Is an instance

    A GWAS analyst intersects a list of disease-associated SNPs against SCREEN, the ENCODE-derived registry of candidate cis-regulatory elements, to flag which variants fall inside an active enhancer rather than an inert intergenic region.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A privately generated ChIP-seq dataset that uses the same assay type as ENCODE's but was never submitted to, quality-checked, or accessioned by the ENCODE Data Coordination Center is not "ENCODE data" -- inclusion requires actual processing through ENCODE's own harmonized pipeline, not just methodological similarity.

Editorial commentary

ENCODE (the Encyclopedia of DNA Elements) is an international research consortium funded by the National Human Genome Research Institute (NHGRI), launched in September 2003 with an explicit mandate to identify every functional element encoded in the human genome — not just the roughly 2% that codes for protein, but the regulatory machinery (promoters, enhancers, transcription-factor binding sites, chromatin states) that controls when and where genes are expressed. NHGRI requires every ENCODE-funded dataset to be released in a free, highly accessible form to all researchers, which is why the project’s output is one of the largest fully open functional-genomics resources in existence.

How ENCODE built its catalog, in phases

ENCODE ran its pilot phase from roughly 2003 to 2007 across about 1% of the genome (30 Mb) to prove the assay strategy would scale. The production phase (2007-2012) extended that strategy genome-wide and generated roughly 15 terabytes of raw sequencing data — the phase that produced the widely cited (and widely debated) 2012 finding that over 80% of the genome shows biochemical activity in at least one assayed cell type, a figure that sparked real scientific disagreement over the distance between “biochemically active” and “functionally important.” Later phases (ENCODE3 and ENCODE4) added depth rather than raw coverage: more cell types and tissues, single-cell and long-read assays, and — new as of ENCODE3 — a parallel mouse genome encyclopedia, so cross-species comparisons of regulatory elements became possible within the same accessioned framework.

What actually gets measured

ENCODE’s core assay set answers different, complementary questions about the same genome: ChIP-seq maps where specific transcription factors and histone modifications physically bind DNA; DNase-seq and ATAC-seq map which regions of chromatin are open and accessible, a strong proxy for regulatory activity; RNA-seq maps what is actually being transcribed, in which cell type; and Hi-C maps which distant genomic regions physically contact each other in 3D, which is how an enhancer far from its target gene in linear sequence gets connected to it. Running the same standardized assay panel across a broad panel of cell lines, primary cells and tissues is what lets a researcher compare, say, a liver-specific enhancer against the same genomic coordinate in a neuron — something a single-lab, single-assay dataset can’t offer.

Data organization and access

Every ENCODE experiment and file carries a standardized accession — ENCSR-prefixed for an experiment (a full submission with its own metadata, protocols, and materials), ENCFF-prefixed for an individual output file. Files are distributed as bigBed or bigWig for track visualization or hic format for 3D-contact data, viewable directly as UCSC-style track hubs. No account or application is required to view or download released data — it is licensed under Creative Commons, and the ENCODE Consortium’s own citation guidance asks only that a publication cite the specific dataset accessions used plus the most recent ENCODE Consortium reference paper, not that a researcher request permission. A central Data Coordination Center is what makes this usable at scale: it enforces consistent ontology-based metadata (cell type, assay, target, biological replicate) across every submitting lab before data is released, which is the actual mechanism behind ENCODE’s reputation for being harmonized rather than merely aggregated.

SCREEN: the derived cis-regulatory registry

Raw and processed ENCODE assay data feeds a separate, purpose-built resource called SCREEN (Search Candidate cis-Regulatory Elements by ENCODE) — a browsable registry of candidate cis-regulatory elements (cCREs) computed by integrating DNase/ATAC accessibility with ChIP-seq signal genome-wide. For a researcher who doesn’t need raw experiment files, SCREEN is often the faster entry point: it answers “is this genomic coordinate a candidate promoter, enhancer, or CTCF-bound element, and in which cell types” directly, without requiring the user to reconstruct that inference from individual ChIP-seq and accessibility tracks themselves.

Where ENCODE fits among named repositories

ENCODE is a functional-annotation resource, not a variant catalog or biobank — it answers “what does this stretch of DNA do, and in which cell type,” a genuinely different question from ClinVar’s variant-pathogenicity claims, gnomAD’s population allele frequencies, or dbSNP’s variant catalog itself. A data management plan citing ENCODE should specify accession-level reuse (ENCSR/ENCFF IDs, or a SCREEN cCRE ID) rather than describing it generically as “public genomics data” — the same accession-level specificity CASRAI’s own Data Management Plan guidance recommends for any named repository dependency.

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="ENCODE (Encyclopedia of DNA Elements)"
      vocab-term-identifier="https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode",
  "name": "ENCODE (Encyclopedia of DNA Elements)",
  "identifier": "https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode",
  "description": "An NHGRI-funded international research consortium, launched in September 2003, that systematically catalogs the functional and regulatory elements of the human genome (and, since ENCODE3, the mouse genome) using standardized assays -- ChIP-seq for transcription-factor binding and histone marks, RNA-seq for transcription, DNase-seq/ATAC-seq for chromatin accessibility, and Hi-C for 3D genome contacts -- run across a wide panel of cell types and tissues. Data is released through the ENCODE Portal under standardized accessions (ENCSR for experiments, ENCFF for files), validated and harmonized by a central Data Coordination Center, and no dataset is 'ENCODE data' unless it was actually submitted to and processed through that pipeline.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
  "url": "https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-09-01T08:01:13",
  "dateModified": "2026-09-01T08:01:13",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →