Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

ProteomeXchange: The Proteomics Repository Consortium and PXD Accession Numbers

ProteomeXchange is the international consortium that coordinates the world’s mass spectrometry proteomics repositories under a single PXD accession-number system. This guide explains how it works, who its members are, what submission requires, and how it maps to journal data-availability policies.

Researchers depositing mass spectrometry (MS)-based proteomics data are often told to “submit to ProteomeXchange,” but ProteomeXchange itself does not store data. It is a consortium — a coordinating framework that links a set of independently operated repositories under one submission standard, one accession-number scheme, and one central search index. Understanding that distinction is the key to understanding how proteomics data sharing actually works, and why a single PXD number can point to a dataset physically hosted at any one of several institutions worldwide.

What ProteomeXchange is — and isn’t

ProteomeXchange (PX) was established to provide globally coordinated standard data submission and dissemination pipelines across the world’s leading MS-based proteomics repositories, and to promote open data practices in the field. It does not run its own storage infrastructure for raw data; instead, member repositories each host data using their own systems, while ProteomeXchange defines shared submission requirements, a common accession-number format, and ProteomeCentral, the central index that aggregates and makes datasets from every member repository searchable in one place.

This consortium model is deliberately similar in spirit to how other domain-specific data infrastructures federate independently operated repositories under shared standards — the same coordination problem that generalist and domain data repositories in other fields solve in different ways. Within proteomics specifically, PX exists because MS instrument output, search-engine results, and experimental metadata are complex and heterogeneous enough that a single common submission standard, rather than each journal or repository inventing its own, was necessary for the field to standardize on public data deposition at scale.

Member repositories

ProteomeXchange is not a single website; it is the standard that a defined set of member repositories submit to and register through. The core members are:

  • PRIDE (PRoteomics IDEntifications database) — hosted at EMBL-EBI (Cambridge, UK); a founding PX member and the largest repository in the consortium by volume.
  • PeptideAtlas — hosted at the Institute for Systems Biology (Seattle, USA); also a founding PX member, built around a compiled, cross-experiment peptide/protein observation resource rather than a simple deposit-and-store model.
  • MassIVE (Mass Spectrometry Interactive Virtual Environment) — hosted at UC San Diego (USA).
  • jPOST (Japan ProteOme STandard Repository) — the Japanese national proteomics repository.
  • iProX — hosted by the National Center for Protein Sciences (Beijing, China).
  • Panorama Public — hosted at the University of Washington (USA); specialized for targeted/quantitative MS data (e.g., SRM/PRM, DIA) built on the Skyline ecosystem.

Each member repository operates under its own governance, funding, and long-term preservation commitments, but all agree to a shared PX Membership Agreement that binds them to the consortium’s common submission requirements and to registering every accepted dataset with ProteomeCentral so it is discoverable regardless of which repository actually stores it. Researchers choosing where to submit generally pick based on geography, data type (e.g., Panorama Public for targeted proteomics, PeptideAtlas for certain reprocessed/compiled datasets), or funder/institutional preference — the choice does not change how the dataset is found or cited, because the PXD number and ProteomeCentral listing work identically across members.

PXD accession numbers

Every original dataset submitted through a ProteomeXchange member repository is assigned a unique PX accession number in the format PXD followed by six digits (e.g., PXD012345). Datasets that represent a reanalysis or reprocessing of previously deposited raw data — rather than a new original submission — are assigned a related but distinct prefix, PXD-style numbers reserved for reanalysis projects (documented in ProteomeXchange’s own reanalysis-dataset guidelines), keeping reprocessed results traceable back to the original submission without overwriting it.

The PXD number functions as a persistent identifier for the dataset: it is the citable reference used in manuscripts, in the dataset’s own metadata record, and by ProteomeCentral’s cross-repository search. A PXD number resolves to a ProteomeCentral record regardless of which member repository actually hosts the underlying files, which is the practical payoff of the consortium model — a reader does not need to know in advance that a given PXD number lives at PRIDE versus MassIVE versus jPOST to find and retrieve it.

Submission requirements: Complete vs. Partial

ProteomeXchange’s data submission guidelines define two broad submission categories, and the distinction matters for anyone planning a deposit around a manuscript deadline:

  • Complete submissions are fully annotated, quality-assured datasets ready for public distribution: raw MS data files, processed/analyzed output (peptide and protein identifications, and quantification results where applicable), and comprehensive experimental metadata covering sample preparation, instrumentation, and analysis parameters.
  • Partial submissions (“supported by repository but incomplete”) are raw or processed data that a member repository has accepted and assigned a PXD number to, but for which the full metadata or final quality-control annotation is not yet complete.

Because reviewers and downstream re-users generally expect Complete status before treating a dataset as fully reusable, the Partial category exists mainly as an interim state — for example, to secure an accession number and a private reviewer-access link before a manuscript’s data curation is finished — rather than as a permanent designation. Submission is handled through the tools each member repository provides (e.g., PRIDE’s submission tool for PRIDE deposits); ProteomeXchange itself does not run a single universal upload portal, but its shared guidelines and metadata schema are what let a Complete submission at any member repository mean the same thing.

Embargo, public release, and journal data-availability requirements

A common workflow is to submit data at the time of manuscript submission but keep it under a reviewer-access embargo, then release it publicly once the paper is accepted or published — mirroring how other domain repositories support data availability statements tied to a publication timeline. ProteomeXchange member repositories support exactly this pattern: a private reviewer link during peer review, followed by public release through ProteomeCentral on acceptance or publication, so the PXD number cited in the paper’s data-availability statement becomes resolvable to the public dataset at (or shortly after) publication.

This matters because proteomics journals increasingly require deposition as a condition of publication rather than a recommendation. Molecular & Cellular Proteomics (MCP), one of the field’s most prominent journals, has required raw MS data deposition with every submitted manuscript since July 2015, and other major publishers (including Nature-family and PLOS journals) have moved toward similar requirements for proteomics submissions. In practice, this means a proteomics manuscript’s data availability statement is frequently just a PXD accession number plus the hosting repository’s name — a pattern that mirrors, at the domain-specific level, the broader push (from ICMJE, funders, and journals generally) toward machine-resolvable, repository-hosted evidence of the underlying data rather than a “data available on request” statement.

Why a consortium model, not a single repository

The consortium structure lets ProteomeXchange do two things a single centralized repository could not do as easily: it allows regionally or institutionally appropriate hosting (a Chinese-funded lab depositing to iProX, a Japanese lab to jPOST, without needing to route data through a foreign-hosted system), and it lets specialized repositories serve specialized data types (Panorama Public’s Skyline-based infrastructure for targeted/quantitative workflows, PeptideAtlas’s cross-study compiled resource) while still presenting one unified, standards-based front end to journals, funders, and re-users through ProteomeCentral. This is the same underlying tension that generalist-versus-domain-specific repository choices reflect elsewhere in research data management — proteomics resolved it by federating domain repositories under a shared identifier and metadata standard rather than picking one winner.

ProteomeXchange has also been designated a Global Core Biodata Resource, a recognition (alongside bodies like GEO/SRA in genomics) reflecting its role as foundational, community-sustained infrastructure that other research depends on for long-term data preservation — the same rationale that underlies federal “desirable characteristics of data repositories” guidance now shaping funder repository-selection policy more broadly.

Practical guidance for researchers and research offices

  • Submit through a member repository, not ProteomeXchange directly — pick PRIDE, MassIVE, jPOST, iProX, PeptideAtlas, or Panorama Public based on data type and institutional/geographic fit; the PXD number and ProteomeCentral listing behave identically afterward.
  • Submit early, embargo if needed — request a PXD accession number and reviewer-access link at manuscript submission rather than after acceptance, since many journals now require the accession number in the submitted manuscript itself.
  • Aim for Complete, not just Partial, status before public release — Partial submissions are a legitimate interim state, but reviewers, funders, and re-users generally expect full metadata and QC annotation for a dataset to be treated as properly deposited.
  • Cite the PXD number, not a repository-specific ID, in the data availability statement — this is what makes the dataset findable through ProteomeCentral regardless of which member repository a future reader checks first.

Frequently asked questions

Is ProteomeXchange itself a data repository?

No. ProteomeXchange is a consortium and coordinating standard; the actual data is hosted by member repositories (PRIDE, PeptideAtlas, MassIVE, jPOST, iProX, Panorama Public). ProteomeCentral, ProteomeXchange’s central index, aggregates records from all of them for unified search, but does not itself store the raw files.

What does a PXD accession number look like?

The prefix “PXD” followed by six digits (for example, PXD012345), assigned to an original dataset submission. Reanalysis/reprocessing datasets built from previously deposited raw data receive a related but distinct accession under ProteomeXchange’s reanalysis-dataset guidelines.

Do I need to deposit to ProteomeXchange to publish in a proteomics journal?

Many proteomics journals now require it. Molecular & Cellular Proteomics has mandated raw MS data deposition with every submission since July 2015, and other major publishers have adopted similar proteomics data-deposition requirements. Check the specific journal’s author guidelines, but treat ProteomeXchange deposition as the default expectation for MS-based proteomics manuscripts, not an optional extra.

Which member repository should I use?

There is no single correct answer — it typically depends on data type (e.g., Panorama Public for targeted/quantitative MS workflows), institutional or funder affiliation, and geography (jPOST for Japan-based work, iProX for China-based work). Because all members share the same PXD accession system and register with ProteomeCentral, the choice does not affect discoverability.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →