Examples
Worked examples
- Is an instance
A cancer biology researcher building a pan-cancer survival model pulls TCGA breast- and lung-adenocarcinoma cohorts from the GDC Data Portal, filtered by a specific somatic mutation, to compare gene-expression patterns across the two cancer types.
- Is an instance
A bioinformatics course assigns TCGA's publicly available, harmonized RNA-seq data as a teaching dataset specifically because it is static, well-documented, and processed through a single reference pipeline -- unlike a live, still-growing repository, results reproduced against a specific TCGA release stay reproducible indefinitely.
Counter-examples
Looks similar, but isn't
- Not an instance
A cancer dataset generated after 2018 by a different NCI initiative -- for example TARGET (pediatric cancers), or a directly-submitted GDC project -- is not TCGA data merely because it is hosted in the same GDC portal; TCGA specifically refers to the original 2006-2018 program's fixed set of 33 cancer types and roughly 11,000 cases.
Editorial commentary
The Cancer Genome Atlas (TCGA) was a landmark, 12-year (2006-2018) cancer genomics program jointly run by the National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI). Over its run, TCGA molecularly characterized more than 11,000 primary cancer cases spanning 33 cancer types, drawing on contributions from thousands of researchers across dozens of institutions. TCGA is now complete — it is a fixed, closed dataset, not an ongoing collection effort.
What kind of data TCGA generated
TCGA’s defining feature was multi-omic characterization of each case, not just DNA sequencing: genomic sequence alterations, epigenomic data (DNA methylation, non-coding RNA), transcriptomic gene-expression measurements, and proteomic (protein expression and structure) data were all generated for the same set of tumor samples wherever feasible. That matched, multi-layered design is what has let the dataset remain heavily reused for pan-cancer and multi-omic integration analyses years after data generation stopped.
Where TCGA data lives now
TCGA does not have its own standalone active portal. Its full dataset was migrated into and is now served through the Genomic Data Commons (GDC), which reprocesses TCGA data through the same harmonized bioinformatics pipeline it applies to other NCI cancer programs (such as TARGET) and directly-submitted projects — meaning a researcher can compare a TCGA cohort against a newer GDC-hosted cohort on equal, identically-processed footing. TCGA data is also widely mirrored and re-indexed by third-party analysis tools, most notably cBioPortal for Cancer Genomics, which many researchers use as their day-to-day interface for browsing TCGA results even though the GDC remains the canonical source.
Examples
- A cancer biology researcher building a pan-cancer survival model pulls TCGA breast- and lung-adenocarcinoma cohorts from the GDC Data Portal, filtered by a specific somatic mutation, to compare gene-expression patterns across the two cancer types.
- A bioinformatics course assigns TCGA’s publicly available, harmonized RNA-seq data as a teaching dataset specifically because it is static, well-documented, and processed through a single reference pipeline — unlike a live, still-growing repository, results reproduced against a specific TCGA release stay reproducible indefinitely.
Counter-example
A cancer dataset generated after 2018 by a different NCI initiative — for example TARGET (pediatric cancers), or a directly-submitted GDC project — is not TCGA data merely because it is hosted in the same GDC portal; TCGA specifically refers to the original 2006-2018 program’s fixed set of 33 cancer types and roughly 11,000 cases.
Related infrastructure
See the Genomic Data Commons (GDC) for how TCGA data is accessed and harmonized today, and dbGaP for the controlled-access mechanics that gate TCGA’s germline/identifiable-risk data tier.
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="The Cancer Genome Atlas (TCGA)"
vocab-term-identifier="https://casrai.org/dictionary/term/the-cancer-genome-atlas-tcga" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/the-cancer-genome-atlas-tcga",
"name": "The Cancer Genome Atlas (TCGA)",
"identifier": "https://casrai.org/dictionary/term/the-cancer-genome-atlas-tcga",
"description": "The Cancer Genome Atlas (TCGA) was a completed (2006-2018), joint National Cancer Institute (NCI) and National Human Genome Research Institute (NHGRI) program that molecularly characterized more than 11,000 primary cancer cases across 33 cancer types, generating matched genomic, epigenomic, transcriptomic, and proteomic data for each case. TCGA is not an active, growing repository -- no new cases have been added since the program concluded -- and its full, fixed dataset now lives inside the Genomic Data Commons (GDC) alongside other harmonized NCI cancer datasets.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
"url": "https://casrai.org/dictionary/term/the-cancer-genome-atlas-tcga",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"author": {
"@id": "https://casrai.org/#editorial-team"
},
"datePublished": "2026-09-01T07:46:15",
"dateModified": "2026-09-01T07:46:15",
"inLanguage": "en-GB",
"isAccessibleForFree": true
}






