Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack BProposedv2026.1

KEGG (Kyoto Encyclopedia of Genes and Genomes)

A manually curated, integrated database resource -- initiated in 1995 by Minoru Kanehisa at Kyoto University's Institute for Chemical Research and maintained by the Kanehisa Laboratories -- that represents biological systems as a computable network linking genes, proteins, small molecules and reactions to the higher-level pathways and functions they participate in. Organized into linked component databases: PATHWAY (manually drawn pathway maps), MODULE (functional gene-set units), BRITE (hierarchical functional classifications), GENES/ORTHOLOGY (gene and ortholog annotation across sequenced genomes), and COMPOUND/REACTION/ENZYME (chemical and enzymatic data). Free to browse via the website; bulk FTP downloads have required a paid subscription since July 2011, a funding-driven access split worth noting explicitly in a data management plan rather than assuming KEGG data is unconditionally free to redistribute.

ByCASRAI Editorial Board
· Last updated 1 Sept 2026
Share this

Ask about KEGG (Kyoto Encyclopedia of Genes and Genomes)

Answers are drawn from this dictionary entry and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Examples

Worked examples

  • Is an instance

    A transcriptomics researcher runs a KEGG pathway enrichment analysis on a list of differentially expressed genes to identify which manually curated metabolic or signaling pathways are statistically overrepresented in the result set.

  • Is an instance

    A comparative-genomics team uses KEGG Orthology (KO) identifiers to map equivalent genes across several sequenced bacterial genomes, since KO groups genes by conserved function rather than by sequence similarity alone.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A bulk mirror of KEGG's pathway maps redistributed without a paid FTP subscription license is not a legitimate reuse of KEGG data as of the 2011 access-model change -- free access covers interactive web browsing, not unrestricted bulk redistribution, and a DMP citing KEGG as a data source should note that distinction rather than assume unconditional free reuse.

Editorial commentary

KEGG (Kyoto Encyclopedia of Genes and Genomes) is an integrated database resource that represents biological systems — cells, organisms, and ecosystems — as a computable network connecting genes and proteins to the molecular reactions, pathways, and higher-level functions they participate in. Minoru Kanehisa initiated KEGG in 1995 at Kyoto University’s Institute for Chemical Research as part of the Japanese Human Genome Program; the Kanehisa Laboratories continue to maintain and expand it. Rather than one flat database, KEGG is organized as several linked component databases that each answer a different question about the same underlying biology.

How KEGG is organized

The PATHWAY database holds manually drawn reference pathway maps representing metabolic, signaling, and disease processes — the resource most researchers mean when they say “KEGG pathway.” MODULE groups genes into smaller functional units below the level of a full pathway. BRITE provides hierarchical functional classifications for genes, proteins, and other biological entities. On the genomic side, GENES and ORTHOLOGY (KO) annotate genes across every sequenced genome KEGG has integrated, with KO groups clustering genes by conserved function across organisms rather than by raw sequence similarity — the property that makes KEGG Orthology useful for comparative genomics. On the chemical side, COMPOUND, GLYCAN, REACTION, and enzyme nomenclature data connect the molecules involved to the reactions and pathways that consume or produce them, plus a DISEASE/DRUG layer connecting pathway-level biology to approved pharmaceuticals and known disease associations.

The access model, and why it matters for a DMP

KEGG remains freely browsable through its website for any researcher. That changed for bulk access in July 2011, when the Kanehisa Laboratories introduced a paid subscription requirement for FTP downloads of the full dataset, citing a significant cutback in government funding — a genuinely disruptive moment for the bioinformatics community that sparked wider discussion about the sustainability of free, government-funded reference databases more broadly. The practical consequence: a researcher who only queries individual pathways or genes interactively through the website is unaffected, but a project planning to mirror, bulk-analyze, or redistribute KEGG data at scale needs to budget for and document a subscription, not assume unconditional free reuse the way many other public bioinformatics resources allow.

What KEGG is used for in practice

The most common research use is pathway enrichment analysis: given a gene list from a differential-expression or GWAS study, testing which KEGG pathways are statistically overrepresented reveals which biological processes the result set implicates, beyond what any single gene’s annotation would show alone. KEGG Orthology identifiers are also widely used as a normalization layer in comparative and metagenomic studies, letting a researcher compare functional gene content across organisms or samples without needing every genome to share direct sequence orthologs.

Where KEGG fits among named repositories

KEGG occupies similar ground to Reactome — both are manually curated pathway resources — but the two differ in curation model and scope: Reactome pathways go through a structured, peer-reviewed editorial process organized around human biology with computational projection to other species, while KEGG’s pathway maps are drawn and maintained centrally by the Kanehisa Laboratories across a broader span of organisms and chemical/disease data. Many analyses run both in parallel and report where the two agree or diverge, rather than treating either as a complete substitute for the other. A DMP should name which resource (and which access tier, given KEGG’s subscription model) an analysis actually depends on, the same accession-level specificity CASRAI’s own Data Management Plan guidance recommends for any named repository dependency.

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="KEGG (Kyoto Encyclopedia of Genes and Genomes)"
      vocab-term-identifier="https://casrai.org/dictionary/term/kyoto-encyclopedia-of-genes-and-genomes-kegg" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/kyoto-encyclopedia-of-genes-and-genomes-kegg",
  "name": "KEGG (Kyoto Encyclopedia of Genes and Genomes)",
  "identifier": "https://casrai.org/dictionary/term/kyoto-encyclopedia-of-genes-and-genomes-kegg",
  "description": "A manually curated, integrated database resource -- initiated in 1995 by Minoru Kanehisa at Kyoto University's Institute for Chemical Research and maintained by the Kanehisa Laboratories -- that represents biological systems as a computable network linking genes, proteins, small molecules and reactions to the higher-level pathways and functions they participate in. Organized into linked component databases: PATHWAY (manually drawn pathway maps), MODULE (functional gene-set units), BRITE (hierarchical functional classifications), GENES/ORTHOLOGY (gene and ortholog annotation across sequenced genomes), and COMPOUND/REACTION/ENZYME (chemical and enzymatic data). Free to browse via the website; bulk FTP downloads have required a paid subscription since July 2011, a funding-driven access split worth noting explicitly in a data management plan rather than assuming KEGG data is unconditionally free to redistribute.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
  "url": "https://casrai.org/dictionary/term/kyoto-encyclopedia-of-genes-and-genomes-kegg",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-09-01T08:05:46",
  "dateModified": "2026-09-01T08:05:46",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →