ArrayExpress is EMBL-EBI’s archive for functional-genomics data — microarray and, later, sequencing-based studies of gene expression and related molecular assays. Established at the European Bioinformatics Institute (EMBL-EBI) in 2002 as the first MIAME-compliant public database, it has served for two decades as the European counterpart to NCBI’s Gene Expression Omnibus (GEO), and journal data-availability policies for gene-expression studies typically name both repositories as acceptable deposit destinations.
This guide covers what ArrayExpress hosts, its 2021 restructuring into the BioStudies database, submission requirements and the Annotare submission tool, how its accession-number system works, and how it relates to GEO.
What ArrayExpress is, and who runs it
ArrayExpress is operated by EMBL-EBI (the European Bioinformatics Institute, part of the European Molecular Biology Laboratory), based in Hinxton, UK. It stores data from high-throughput functional-genomics experiments, including detailed sample annotations, experimental protocols, and both processed and raw data. Like GEO, it was built around the MIAME (Minimum Information About a Microarray Experiment) standard developed by the MGED Society (now the Functional Genomics Data, or FGED, Society), and it has since extended to cover sequencing-based functional-genomics studies as well.
A structural detail worth knowing: for high-throughput sequencing studies, ArrayExpress does not store raw sequence reads itself. Those are brokered to the European Nucleotide Archive (ENA), EMBL-EBI’s own sequence-data repository, while ArrayExpress retains the processed data and the sample/experimental metadata linking back to the raw reads. This is a similar division of labour to how GEO relies on NCBI’s Sequence Read Archive (SRA) for raw reads.
What data types ArrayExpress hosts
ArrayExpress’s scope covers functional-genomics assay types broadly, not just classic expression microarrays:
- Microarray-based gene-expression studies (its original, legacy scope)
- RNA-seq and other high-throughput sequencing-based expression studies
- Single-cell RNA-seq data
- ChIP-seq, methylation, and other functional-genomics assay types
- Associated metadata: sample annotations, experimental protocols, and processed data matrices
As with GEO, it is a discipline-specific repository for functional genomics, not a general-purpose research-data archive — see CASRAI’s overview of discipline-specific repositories versus generalist repositories, and the broader guide to research repositories versus archives, for data that falls outside functional genomics.
The 2021 migration into BioStudies
In 2021, EMBL-EBI migrated ArrayExpress’s holdings into BioStudies, a database established in 2016 to handle multimodal biological data — studies that combine several assay types (for example, transcriptomic, proteomic, and genotyping data from the same experiment) rather than a single data modality. EMBL-EBI’s rationale, described in the group’s own Nucleic Acids Research paper on the migration, was that functional-genomics experiments were increasingly multimodal in a way the original ArrayExpress database structure wasn’t designed to accommodate.
Practically, this migration was designed to be seamless for depositors and users: existing accession numbers were retained, the submission tool (Annotare) stayed the same, and the ArrayExpress collection remains browsable within BioStudies at the same conceptual location, now presented as the “ArrayExpress collection” of BioStudies. Anyone citing or searching for an ArrayExpress accession today will find it served from BioStudies infrastructure rather than a standalone ArrayExpress database, but the accession number itself, and the deposit workflow, are unchanged.
Submission requirements
Data is submitted to ArrayExpress through Annotare, EMBL-EBI’s web-based submission tool, which walks depositors through providing:
- Experimental design and protocol descriptions sufficient to satisfy MIAME (for array-based studies) or the equivalent minimum-information expectations for sequencing-based studies
- Sample-level metadata (organism, tissue/cell type, treatment conditions, replicate structure)
- Processed data files, and — for sequencing studies — a link to the corresponding raw-read submission in the European Nucleotide Archive (ENA), since ArrayExpress itself does not host raw sequence reads
As with GEO, submitted studies receive a permanent, citable accession number and can be kept private for a limited embargo period, becoming public either when the accession is cited in a publication or at a submitter-specified release date — the same publish-on-citation-or-date pattern most major domain repositories use to accommodate pre-publication peer review.
Accession-number conventions
ArrayExpress accessions follow an E-xxxx-nnnn pattern, where the middle code indicates the submission route or data source:
- E-MTAB-nnnn — the current, standard prefix for studies submitted in MAGE-TAB format via Annotare; this is what the large majority of active ArrayExpress accessions use today.
- E-GEOD-nnnn — denotes a study that was imported into ArrayExpress from NCBI’s GEO, rather than submitted directly.
- E-MEXP-nnnn — a legacy prefix from MIAMExpress, an earlier submission route retired in 2014; these are older studies still resolvable in the archive but not an active submission path.
The prefix is a useful quick signal when browsing: an E-GEOD accession tells you the record originated as a GEO deposit and was mirrored into ArrayExpress, while an E-MTAB accession was submitted natively.
How ArrayExpress relates to GEO
ArrayExpress and GEO are independently governed — one run by EMBL-EBI in the UK/EU, the other by NCBI (part of the U.S. National Institutes of Health) — but they serve the same research community and were both built around the same MIAME/MINSEQE minimum-information standards, which is why journal and funder data-availability policies for gene-expression studies typically list them interchangeably as acceptable deposit destinations. See CASRAI’s GEO guide for the equivalent detail on NCBI’s side.
The two archives have a documented history of data exchange: the E-GEOD accession prefix exists specifically to mark studies imported into ArrayExpress from GEO. This reflects a historical import mechanism rather than a live, automatic two-way sync for every deposit — a researcher choosing where to submit should not assume that depositing in one repository automatically creates a citable accession in the other. When a study needs to be discoverable through both ecosystems, the safest approach is to check the specific journal’s or funder’s data-availability policy and, where dual deposit is expected or preferred, submit to both repositories directly rather than relying on cross-import.
In practice, the choice between ArrayExpress and GEO for a new submission is often driven by geography and community convention (European Molecular Biology Organization-affiliated and many European-funded projects default to ArrayExpress/BioStudies; many others default to GEO), by which raw-sequence archive a lab is already using (ENA versus SRA), or simply by a specific journal’s stated preference — rather than by any functional difference in what the two repositories can accept, since both are built around the same underlying minimum-information standards.
Frequently asked questions
Is ArrayExpress still active, or has it been replaced?
ArrayExpress is still active as a data collection, but as of 2021 it operates as part of EMBL-EBI’s BioStudies database rather than as a standalone system. Existing and new ArrayExpress accessions, and the Annotare submission workflow, are unaffected by the change.
What is Annotare?
Annotare is EMBL-EBI’s web-based tool for submitting functional-genomics studies to ArrayExpress. It collects experimental design, protocol, and sample metadata alongside the processed data files.
What does an E-MTAB or E-GEOD accession number mean?
The prefix identifies the submission route: E-MTAB denotes a study submitted directly via Annotare in MAGE-TAB format (the current standard route), E-GEOD denotes a study imported from NCBI’s GEO, and E-MEXP denotes a legacy submission via the now-retired MIAMExpress tool.
Do I need to submit my data to both ArrayExpress and GEO?
Not automatically — the two repositories do not perform a live two-way sync for every new deposit. Check your target journal’s or funder’s data-availability policy; where both are acceptable, choose one based on institutional/community convention, or submit to both directly if dual discoverability is required.
Does ArrayExpress store raw sequencing reads?
No. For sequencing-based studies, raw reads are brokered to the European Nucleotide Archive (ENA), EMBL-EBI’s sequence-data repository; ArrayExpress holds the processed data and the metadata linking to the corresponding ENA raw-read submission.







