Metabolomics Workbench is a public data repository and set of analysis tools for metabolomics research, funded by the NIH Common Fund’s Metabolomics Program and operated as the U.S. National Metabolomics Data Repository (NMDR). For a researcher deciding where to deposit raw and processed metabolomics data, or a research administrator reviewing a data management plan (DMP) that names it as the repository of record, this guide covers what the platform actually is, who runs it, what it requires at submission, and how its identifier system works.
What Is Metabolomics Workbench?
Metabolomics Workbench describes itself as a national and international repository for metabolomics data and metadata, alongside metabolite standards, protocols, tutorials, training materials, and analysis tools. It is not simply a file-dump archive: every deposited study is expected to carry structured experimental metadata alongside the raw and processed measurements, so that a dataset remains interpretable and reusable once it leaves the hands of the lab that produced it. This positions it as a domain repository in the same sense as Gene Expression Omnibus (GEO) for transcriptomics or ProteomeXchange for proteomics: a discipline-specific repository built around the metadata conventions and file formats a particular measurement community actually produces, rather than a generalist archive.
Who Runs It: NIH Common Fund Metabolomics Program and NMDR
The National Metabolomics Data Repository is housed at the San Diego Supercomputer Center (SDSC), University of California, San Diego, and receives support through NIH Common Fund Metabolomics Program grants, including contributions from the Common Fund Data Ecosystem. The program’s broader mandate is to increase national capacity in metabolomics: supporting technology development, training, reference-standard availability, and data sharing. Metabolomics Workbench is the data-repository component of that mandate, functioning as the field’s designated place of deposit in the U.S., comparable in role to how NCBI’s GEO functions for gene expression data or the NIH’s dbGaP functions for genotype-phenotype data (see CASRAI’s dbGaP entry for that comparison).
What Data Types It Hosts
The repository covers both major metabolomics measurement modalities: mass spectrometry (MS) and nuclear magnetic resonance (NMR). Within MS, it accepts data from a wide range of analytical platforms and ionization/separation techniques, including LC-MS, GC-MS, CE-MS, MALDI-MS, direct-infusion MS, APCI-MS, API-MS, DESI-MS, ICR-MS, and flow-injection analysis MS (FIA-MS). Studies span a broad range of organisms, human, mouse, and other model organisms such as Drosophila melanogaster, along with fish species and bacterial samples, and a wide range of sample sources: tissues (adipose, brain), biological fluids (plasma, serum, urine, feces), and cultured cells. Alongside the deposited study data, the platform hosts a metabolite structure database, spectral libraries, and reference standards, extending its role beyond a passive archive into a working analysis resource.
Submission Requirements and the mwTab Metadata Standard
Metabolomics Workbench requires that experimental metadata be deposited alongside the metabolite measurements themselves, on the reasoning that measurements without full experimental context (sample preparation, instrument parameters, study design) cannot be reliably reproduced or reused. The repository’s submission tool guides a depositor through three stages: registering the study, submitting metadata and processed data, and uploading raw instrument data and supplementary material. The standard interchange format for this is mwTab, a flat-file format developed specifically by Metabolomics Workbench to bundle metadata and processed results in one file, and which every deposited study is stored and served as, alongside a JSON equivalent for programmatic access. A companion open-source tool, the mwtab Python library, was built to support RESTful access, quality control, and deposition/curation against this format. Raw instrument data is generally accepted in vendor-native formats (for example, Thermo raw files) as well as open interchange formats such as mzML and mzXML. Because required fields and file-format expectations are revised periodically, anyone preparing a submission should work from the current tutorial and templates on the Metabolomics Workbench site itself rather than from a fixed checklist, since the exact field list has changed across format versions.
The Study ID and Analysis ID System
Metabolomics Workbench identifies deposited work at two levels. A Study ID, in the format ST followed by six digits (for example, ST004769), identifies the overall study record, its design, participants or samples, and metadata. Because a single study can include more than one analytical run or method, each individual analysis within a study, a specific MS or NMR dataset with its own results, is assigned a separate Analysis ID in the format AN followed by digits. The Analysis ID is the unique, citable identifier for a specific dataset within a study; the Study ID groups related analyses under one experimental record. When citing Metabolomics Workbench data in a manuscript, DMP, or repository cross-reference, cite the specific Analysis ID(s) used, not just the parent Study ID, since a study can bundle multiple distinct analyses with different scope.
Metabolomics Workbench vs. MetaboLights
Metabolomics Workbench’s counterpart in Europe is MetaboLights, hosted by the European Bioinformatics Institute (EBI) as part of ELIXIR, Europe’s distributed life-science data infrastructure. The two repositories are frequently described in the metabolomics-informatics literature as complementary, geographically distinct counterparts, roughly analogous to how the NCBI Sequence Read Archive and EBI’s European Nucleotide Archive both serve as points of deposit for sequencing data on either side of the Atlantic. Metabolomics Workbench and MetaboLights have collaborated on shared or interoperable data-exchange approaches, including work toward common flat-file formats that would let a dataset move between the two systems, though the two repositories use their own distinct submission formats and metadata schemas (mwTab for Metabolomics Workbench; an ISA-Tab-derived format for MetaboLights) rather than a single shared one. For a researcher or research administrator, the practical implication is straightforward: check the specific journal’s or funder’s data-deposition policy before choosing between them, some journals and funders accept either repository, some name one specifically, and a small number of studies are deposited in both to satisfy overlapping requirements. CASRAI covers MetaboLights in its own right separately; this guide is limited to Metabolomics Workbench.
Why This Matters for Research Administration
For a research administrator reviewing a data management plan or NIH Data Management and Sharing Policy (DMSP) submission that names Metabolomics Workbench as the repository of record, the questions worth checking are the same ones that apply to any named domain repository: does the study’s data type actually fall within the repository’s accepted scope (MS or NMR metabolomics data, not, for example, raw genomic sequence); does the plan account for the metadata burden of mwTab submission, since experimental metadata is a hard requirement, not optional context; and is the named Study or Analysis ID (once assigned) recorded in the award’s closeout documentation as evidence the data-sharing commitment was actually met. Unlike institutional or generalist repositories, Metabolomics Workbench does not independently hold a CoreTrustSeal certification listing at the time of writing; its trustworthiness case rests instead on its funding line (NIH Common Fund), long operating history, and community adoption as the field’s de facto data-of-record for U.S.-funded metabolomics work rather than a third-party certification. Administrators evaluating repository choice more generally can use CASRAI’s CoreTrustSeal certification guide or the Desirable Characteristics of Data Repositories checklist as a general evaluation framework.
Frequently Asked Questions
Is Metabolomics Workbench free to use?
Yes. Deposit and access are free; the repository is publicly funded through the NIH Common Fund rather than operating on a subscription or fee-for-deposit model.
What file format does Metabolomics Workbench use?
Its native format is mwTab, a flat text file combining structured experimental metadata with processed metabolite measurement data; each study/analysis is also available as an equivalent JSON file for programmatic use. Raw instrument files are stored separately, typically in vendor-native or open MS interchange formats such as mzML.
How is Metabolomics Workbench different from GEO or ProteomeXchange?
All three are NIH- or community-backed domain repositories that share the same basic model, structured metadata plus measurement data, unique accession identifiers, free public access, but each is built around a different molecular data type and its field-specific standards: GEO for gene expression and other functional genomics data (MIAME/MINSEQE standards), ProteomeXchange for mass-spectrometry proteomics (PXD accessions), and Metabolomics Workbench for MS/NMR metabolomics (mwTab, ST/AN accessions).
Do I need to deposit in both Metabolomics Workbench and MetaboLights?
Not by default. Most studies deposit in one or the other, chosen based on the depositing researcher’s location, funder requirements, or the target journal’s data-availability policy. Deposit in both only if a specific journal or funder requirement explicitly calls for it.







