Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Gene Expression Omnibus (GEO): NCBI’s Repository, MIAME/MINSEQE, and Accession Numbers

GEO is NCBI’s public repository for gene-expression data, built around the MIAME/MINSEQE minimum-information standards and a GSE/GSM/GPL/GDS accession system.

The Gene Expression Omnibus (GEO) is the National Center for Biotechnology Information’s (NCBI) public repository for gene-expression and other high-throughput functional genomics data — microarray, RNA-seq, and related next-generation sequencing datasets. NCBI is itself part of the U.S. National Library of Medicine (NLM), one of the National Institutes of Health (NIH). GEO has been the field’s default deposit destination for transcriptomics data since its launch in 2000, and it is the repository most gene-expression journal data-availability policies are written around, alongside its European counterpart, ArrayExpress.

This guide covers what GEO is and who runs it, the MIAME and MINSEQE submission standards its deposit process is built to satisfy, how its accession-number system (GSE, GSM, GPL, GDS) works, how GEO deposit fits into journal data-availability requirements for gene-expression studies, and how GEO relates to ArrayExpress.

What GEO is, and who runs it

GEO is described in NCBI’s own documentation as “an international public repository that archives and freely distributes microarray, next-generation sequencing, and other forms of high-throughput functional genomics data” submitted by the research community. It is free to use for both submitters and data consumers, and every dataset that clears GEO’s basic curation checks receives a permanent, citable accession number (see below) that resolves to a stable landing page.

GEO’s scope has grown considerably since its 2000 launch: it started as a microarray-era repository and now accepts a broad range of functional genomics data types, including RNA-seq, ChIP-seq, single-cell RNA-seq, methylation arrays, and other high-throughput assay outputs, alongside legacy microarray series. It is not a general-purpose data repository — for research data outside functional genomics, see CASRAI’s overview of research repositories versus archives and the broader distinction between discipline-specific and generalist repositories.

MIAME and MINSEQE: the submission standards GEO is built around

GEO does not simply accept a raw data file. Its submission process is explicitly designed to collect the information specified by two community-developed minimum-information standards, originally developed by the Microarray Gene Expression Data (MGED) Society, now the Functional Genomics Data (FGED) Society:

  • MIAME (Minimum Information About a Microarray Experiment) — the original standard, published in 2001, specifying the information a microarray-based gene-expression study needs to report for the experiment to be independently interpreted and, in principle, reproduced.
  • MINSEQE (Minimum Information about a high-throughput SEQuencing Experiment) — the sequencing-era counterpart to MIAME, extending the same minimum-information philosophy to RNA-seq and other next-generation sequencing-based functional genomics experiments.

Per NCBI’s own GEO documentation, MIAME/MINSEQE compliance “is not related to the submission format or route, but rather to the content provided” — in practice, a submission is compliant if it supplies six categories of information:

  1. Raw data — the underlying assay files for each sample (for example, CEL files for microarrays or FASTQ files for sequencing).
  2. Processed (final) data — the normalized data actually used to support the study’s conclusions, such as a processed gene-expression matrix.
  3. Sample annotation — the biological and experimental variables for each sample: tissue or cell type, organism, sex, age, treatment/compound and dose, and other experimental factors.
  4. Experimental design — how raw data files map to samples, and which samples are technical versus biological replicates of one another.
  5. Feature (platform) annotation — what each measured feature on the array or in the sequencing output actually represents (gene identifiers, probe sequences, genomic coordinates).
  6. Protocols — the laboratory and data-processing methods used, including normalization procedures.

Because GEO’s submission templates (spreadsheet or web-form based) are built to collect exactly these six categories, a researcher who completes a GEO submission in full has, as a byproduct, produced a MIAME/MINSEQE-compliant record — which is the detail that matters most for satisfying a journal or funder policy that references either standard by name.

The GEO accession-number system: GSE, GSM, GPL, GDS

GEO organizes deposited data into four record types, each with its own accession prefix. Understanding the distinction matters for citing data correctly and for locating the right level of granularity when reusing someone else’s deposit:

  • GPL — Platform. Describes the array design or sequencer/technology used to generate the data (for example, a specific microarray chip or sequencing instrument type), including the feature/probe annotation table. Multiple series can reference the same platform record.
  • GSM — Sample. A single biological sample’s record: the experimental conditions under which it was handled, and its measurement results (raw and/or processed values for that one sample).
  • GSE — Series. The record for a complete study: a collection of related GSM sample records that together describe one coherent experiment. A GSE accession is the one most commonly cited in a paper’s data-availability statement, since it represents the whole submitted dataset rather than one sample or one platform.
  • GDS — DataSet. A curated record that NCBI staff assemble from one or more Series, grouping biologically comparable samples so the data can be used directly in GEO’s own analysis tools (such as GEO2R). GDS records exist for only a subset of deposited Series — not every GSE has a corresponding curated GDS — so a GSE accession, not a GDS accession, is the reliable identifier to expect when checking whether a given study’s data is in GEO at all.

In practice, a data-availability statement that says a dataset is “available in GEO under accession GSE123456” is pointing to the Series record, which in turn links out to its constituent GSM sample records and the GPL platform record(s) they were generated on. When evaluating or reusing a GEO deposit, start from the GSE accession and drill down from there.

How GEO deposit fits into journal data-availability requirements

Many journals that publish gene-expression research — this has been standard practice across much of molecular biology and genomics publishing since the early-to-mid 2000s — require, as a condition of publication, that microarray or sequencing data underlying a paper be deposited in a MIAME/MINSEQE-compliant public repository and that the resulting accession number be cited in the paper’s data availability statement. GEO (or ArrayExpress, see below) is the standard way authors satisfy that requirement for gene-expression data specifically, in the same way dbGaP is the standard destination for controlled-access human genotype-phenotype data, or GenBank is for nucleotide sequence data. For the mechanics of writing that statement and citing a repository accession correctly, see CASRAI’s guides to writing a data availability statement and a worked data-availability-statement example.

Funder data-sharing policies interact with this the same way: where a specific NIH-designated or discipline-specific repository exists for a data type, NIH’s own repository-selection guidance places it ahead of a generalist repository — GEO is that discipline-specific repository for functional genomics data, which is why an NIH Data Management and Sharing Plan covering gene-expression data should generally name GEO (or ArrayExpress) rather than a generalist option. CASRAI’s guide to GREI, NIH’s Generalist Repository Ecosystem Initiative, covers where generalist repositories fit in that same hierarchy when no discipline-specific repository exists — GEO is precisely the kind of designated, discipline-specific option that takes priority over a GREI-member generalist repository when the data type is gene expression.

GEO and ArrayExpress

ArrayExpress, operated by EMBL-EBI (the European Bioinformatics Institute) in the UK, is the European equivalent of GEO: a public, MIAME/MINSEQE-oriented repository for functional genomics data, built and governed independently of NCBI. GEO and ArrayExpress are not the same database, but they serve the same functional role for most authors and reviewers — a journal or funder policy that requires deposit in “a MIAME-compliant public repository” is generally satisfied by either one, and researchers typically choose based on regional convention, institutional practice, or which repository a collaborating lab already uses. The two repositories have historically maintained data-exchange arrangements so that data submitted to one is more discoverable from the other, though the exact technical mechanics of that exchange have changed over time. This guide covers GEO specifically; CASRAI does not currently have a dedicated ArrayExpress page.

Frequently asked questions

What does GSE mean in GEO?

GSE is GEO’s accession prefix for a Series record — the complete dataset for one study, made up of one or more GSM sample records. It is the accession most commonly cited in a data-availability statement.

What is the difference between GSE, GSM, GPL, and GDS?

GSE identifies a whole study (a Series of samples), GSM identifies one individual sample within that series, GPL identifies the platform/technology the samples were measured on, and GDS identifies a curated dataset NCBI staff have assembled from one or more series for use in GEO’s own analysis tools. Not every GSE has a corresponding GDS.

Is depositing data in GEO required by journals?

Many journals that publish gene-expression research require deposit of the underlying data in a MIAME/MINSEQE-compliant public repository (GEO or ArrayExpress) as a condition of publication, and require the resulting accession number to appear in the paper’s data availability statement. Requirements vary by journal and publisher, so authors should check the specific journal’s data policy.

What is MINSEQE compliance?

MINSEQE (Minimum Information about a high-throughput SEQuencing Experiment) is the sequencing-era counterpart to MIAME, specifying the minimum information — raw data, processed data, sample and experimental-design annotation, platform/feature annotation, and protocols — needed to interpret a next-generation sequencing-based functional genomics experiment. GEO’s submission process is designed to collect this information directly.

Is GEO the same as ArrayExpress?

No. GEO is operated by NCBI (part of NIH/NLM, United States); ArrayExpress is operated by EMBL-EBI (United Kingdom). They are independently governed repositories that serve the same functional role and both target MIAME/MINSEQE compliance, so most journal and funder policies that require “a public, MIAME-compliant repository” accept a deposit in either one.

How do I find a study’s data using its GEO accession number?

Search the accession number (for example, a GSE number) directly at the GEO website or through NCBI’s search interface. A GSE Series page links out to its constituent GSM sample records and GPL platform record(s), so starting from the Series accession is the most reliable way to locate the complete dataset.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →