Examples
Worked examples
- Is an instance
A regulatory-genomics researcher downloads a bigWig ChIP-seq signal track for a specific transcription factor in a liver cell line from the ENCODE Portal, citing the file's ENCFF accession in the manuscript's data-availability statement.
- Is an instance
A GWAS analyst intersects a list of disease-associated SNPs against SCREEN, the ENCODE-derived registry of candidate cis-regulatory elements, to flag which variants fall inside an active enhancer rather than an inert intergenic region.
Counter-examples
Looks similar, but isn't
- Not an instance
A privately generated ChIP-seq dataset that uses the same assay type as ENCODE's but was never submitted to, quality-checked, or accessioned by the ENCODE Data Coordination Center is not "ENCODE data" -- inclusion requires actual processing through ENCODE's own harmonized pipeline, not just methodological similarity.
Editorial commentary
ENCODE (the Encyclopedia of DNA Elements) is an international research consortium funded by the National Human Genome Research Institute (NHGRI), launched in September 2003 with an explicit mandate to identify every functional element encoded in the human genome — not just the roughly 2% that codes for protein, but the regulatory machinery (promoters, enhancers, transcription-factor binding sites, chromatin states) that controls when and where genes are expressed. NHGRI requires every ENCODE-funded dataset to be released in a free, highly accessible form to all researchers, which is why the project’s output is one of the largest fully open functional-genomics resources in existence.
How ENCODE built its catalog, in phases
ENCODE ran its pilot phase from roughly 2003 to 2007 across about 1% of the genome (30 Mb) to prove the assay strategy would scale. The production phase (2007-2012) extended that strategy genome-wide and generated roughly 15 terabytes of raw sequencing data — the phase that produced the widely cited (and widely debated) 2012 finding that over 80% of the genome shows biochemical activity in at least one assayed cell type, a figure that sparked real scientific disagreement over the distance between “biochemically active” and “functionally important.” Later phases (ENCODE3 and ENCODE4) added depth rather than raw coverage: more cell types and tissues, single-cell and long-read assays, and — new as of ENCODE3 — a parallel mouse genome encyclopedia, so cross-species comparisons of regulatory elements became possible within the same accessioned framework.
What actually gets measured
ENCODE’s core assay set answers different, complementary questions about the same genome: ChIP-seq maps where specific transcription factors and histone modifications physically bind DNA; DNase-seq and ATAC-seq map which regions of chromatin are open and accessible, a strong proxy for regulatory activity; RNA-seq maps what is actually being transcribed, in which cell type; and Hi-C maps which distant genomic regions physically contact each other in 3D, which is how an enhancer far from its target gene in linear sequence gets connected to it. Running the same standardized assay panel across a broad panel of cell lines, primary cells and tissues is what lets a researcher compare, say, a liver-specific enhancer against the same genomic coordinate in a neuron — something a single-lab, single-assay dataset can’t offer.
Data organization and access
Every ENCODE experiment and file carries a standardized accession — ENCSR-prefixed for an experiment (a full submission with its own metadata, protocols, and materials), ENCFF-prefixed for an individual output file. Files are distributed as bigBed or bigWig for track visualization or hic format for 3D-contact data, viewable directly as UCSC-style track hubs. No account or application is required to view or download released data — it is licensed under Creative Commons, and the ENCODE Consortium’s own citation guidance asks only that a publication cite the specific dataset accessions used plus the most recent ENCODE Consortium reference paper, not that a researcher request permission. A central Data Coordination Center is what makes this usable at scale: it enforces consistent ontology-based metadata (cell type, assay, target, biological replicate) across every submitting lab before data is released, which is the actual mechanism behind ENCODE’s reputation for being harmonized rather than merely aggregated.
SCREEN: the derived cis-regulatory registry
Raw and processed ENCODE assay data feeds a separate, purpose-built resource called SCREEN (Search Candidate cis-Regulatory Elements by ENCODE) — a browsable registry of candidate cis-regulatory elements (cCREs) computed by integrating DNase/ATAC accessibility with ChIP-seq signal genome-wide. For a researcher who doesn’t need raw experiment files, SCREEN is often the faster entry point: it answers “is this genomic coordinate a candidate promoter, enhancer, or CTCF-bound element, and in which cell types” directly, without requiring the user to reconstruct that inference from individual ChIP-seq and accessibility tracks themselves.
Where ENCODE fits among named repositories
ENCODE is a functional-annotation resource, not a variant catalog or biobank — it answers “what does this stretch of DNA do, and in which cell type,” a genuinely different question from ClinVar’s variant-pathogenicity claims, gnomAD’s population allele frequencies, or dbSNP’s variant catalog itself. A data management plan citing ENCODE should specify accession-level reuse (ENCSR/ENCFF IDs, or a SCREEN cCRE ID) rather than describing it generically as “public genomics data” — the same accession-level specificity CASRAI’s own Data Management Plan guidance recommends for any named repository dependency.
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="ENCODE (Encyclopedia of DNA Elements)"
vocab-term-identifier="https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode",
"name": "ENCODE (Encyclopedia of DNA Elements)",
"identifier": "https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode",
"description": "An NHGRI-funded international research consortium, launched in September 2003, that systematically catalogs the functional and regulatory elements of the human genome (and, since ENCODE3, the mouse genome) using standardized assays -- ChIP-seq for transcription-factor binding and histone marks, RNA-seq for transcription, DNase-seq/ATAC-seq for chromatin accessibility, and Hi-C for 3D genome contacts -- run across a wide panel of cell types and tissues. Data is released through the ENCODE Portal under standardized accessions (ENCSR for experiments, ENCFF for files), validated and harmonized by a central Data Coordination Center, and no dataset is 'ENCODE data' unless it was actually submitted to and processed through that pipeline.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
"url": "https://casrai.org/dictionary/term/encyclopedia-of-dna-elements-encode",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"author": {
"@id": "https://casrai.org/#editorial-team"
},
"datePublished": "2026-09-01T08:01:13",
"dateModified": "2026-09-01T08:01:13",
"inLanguage": "en-GB",
"isAccessibleForFree": true
}






