The Protein Data Bank (PDB) is the single global archive for experimentally determined 3D structures of proteins, nucleic acids, and their complex assemblies. Since 1971 it has served as the reference repository that structural biology, biochemistry, and structure-based drug discovery depend on: when a paper reports a crystal structure, an NMR ensemble, or a cryo-EM model, the underlying atomic coordinates are almost always deposited in, and made permanently retrievable from, the PDB. For a research administrator, the PDB is worth understanding on its own terms — not just as “a repository,” but as a domain-specific repository with its own governance body, deposition mandates, and data standards that plug directly into funder data management plan requirements.
What the PDB Archives, Specifically
Each PDB entry is a complete structural record, not just a set of coordinates. A typical entry bundles: the atomic coordinates of the molecule (a protein, nucleic acid, protein–nucleic acid complex, or larger assembly); experimental metadata describing how the structure was determined — X-ray crystallography, solution or solid-state NMR, or cryo-electron microscopy, among other methods; and a structured validation report assessing the model’s quality against community-agreed metrics for that experimental technique. This combination — model plus method plus validation — is what distinguishes the PDB from a generic file-hosting repository: every entry has passed through a curation and biocuration process before it is released.
The PDB does not archive raw experimental data (diffraction images, raw cryo-EM micrographs, or NMR free-induction decay data) — those live in separate, purpose-built archives that feed into PDB depositions. Raw cryo-EM image data, for example, belongs in EMPIAR, while the derived 3D density maps belong in the Electron Microscopy Data Bank (EMDB); the PDB holds the final, interpreted atomic model built from that map.
Governance: The Worldwide Protein Data Bank (wwPDB)
The PDB is not run by a single institution. It is jointly governed by the Worldwide Protein Data Bank (wwPDB), a consortium of core member organizations that together operate one unified, non-duplicated global archive rather than separate national copies. As of this writing, wwPDB’s core members are:
- RCSB PDB (Research Collaboratory for Structural Bioinformatics) — the US data center, which biocurates depositions originating from the Americas and provides the widely used rcsb.org search, API, and visualization tools;
- PDBe (Protein Data Bank in Europe), hosted at EMBL-EBI;
- PDBj (Protein Data Bank Japan);
- BMRB (Biological Magnetic Resonance Data Bank), which handles NMR experimental data specifically; and
- EMDB (Electron Microscopy Data Bank), which archives the 3D density maps that underlie cryo-EM PDB entries.
Protein Data Bank China (PDB-CN) joined wwPDB as an associate member in 2022. Structural genomics and depositor-facing work is distributed by geography and technique across these centers, but a structure deposited anywhere in the consortium is biocurated to shared standards and released into the same single global archive — there is only one PDB, mirrored and jointly maintained, not several competing national databases. wwPDB describes its mission as managing the PDB archive as a public good and providing deposition, validation, biocuration, and remediation services free of charge, with the goal of universal open access to public-domain structural biology data. RCSB PDB’s own operations are supported by US federal funding, including the National Science Foundation, the Department of Energy, and the National Institutes of Health (through NIGMS, NCI, and NIAID); the wwPDB partners in Europe and Japan are similarly funded through their respective national and regional science agencies.
wwPDB also develops and maintains the shared data standards, validation criteria, and biocuration procedures the whole consortium applies — work that includes standing Validation Task Forces for X-ray, NMR, and 3D electron microscopy methods, whose recommendations define what a wwPDB-compliant validation report must check for each experimental technique.
Deposition: Why It Happens and What It Requires
PDB deposition is not merely encouraged — for most structural biology publications it is a practical precondition of publication. Many journals that publish structural results (across crystallography, structural biology, and general life-science titles) require authors to deposit coordinates in the PDB and to release an accession code at, or before, publication, mirroring the long-standing norm in genomics of mandatory sequence deposition in GenBank/ENA/DDBJ. Funder data management and sharing policies reinforce the same expectation: US federal funders that support structural biology research generally expect structural coordinates and associated experimental metadata to be deposited in a recognized public repository such as the PDB as part of standard data-sharing obligations, rather than retained only as journal supplementary files.
The practical deposition workflow, coordinated through wwPDB’s OneDep system, involves:
- Structure and metadata submission — coordinates plus experimental details (method, resolution, refinement statistics, sample composition) submitted through a single unified interface regardless of which wwPDB center ultimately processes the entry;
- Biocuration — wwPDB staff check the deposition for consistency, completeness, and correct annotation before release;
- Validation — an automated validation report is generated against the relevant Validation Task Force criteria, flagging geometry, fit-to-data, and other quality issues the depositor can address before release;
- Release — entries are typically released immediately, on publication, or after a depositor-specified hold (commonly up to one year), after which point they become permanently, freely accessible.
Every released entry receives a unique four-character PDB ID (a persistent identifier in practice, though not built on the DOI infrastructure used elsewhere in research data management) that is cited in the literature and never reassigned, which is what makes a PDB structure reliably re-findable and reusable years after deposition.
Data Formats: Legacy PDB Format and PDBx/mmCIF
The PDB’s original flat-file format (fixed-column “PDB format”) is still widely recognized by structural biology software, but it has real structural limitations — a rigid column layout that cannot cleanly represent very large assemblies, certain modern experimental metadata, or some chemical component definitions. wwPDB’s current master format is PDBx/mmCIF (macromolecular Crystallographic Information File), a more extensible, dictionary-defined format capable of representing structures and metadata the legacy format cannot. Since wwPDB moved to PDBx/mmCIF as its master format, legacy-format files for many entries are generated as a derived, best-effort conversion rather than the authoritative source — researchers building tooling against the PDB archive should treat PDBx/mmCIF, not the legacy format, as the canonical representation going forward.
The PDB Within the Structural Biology Archive Ecosystem
Researchers depositing or reusing structural data need to place the PDB correctly relative to its sibling archives, since each holds a different stage of the same underlying workflow:
- Raw image/experimental data (e.g., cryo-EM micrographs, tilt series) — archived in EMPIAR, not the PDB.
- Derived 3D density maps (cryo-EM reconstructions before atomic interpretation) — archived in EMDB.
- Final, interpreted atomic models — archived in the PDB itself, cross-referenced back to the EMDB map or crystallographic/NMR data that produced them.
This three-tier separation exists because each stage has a genuinely different reuse profile: raw images are needed mainly by methods developers reprocessing data with newer algorithms, density maps are needed for structure validation and re-refinement, and atomic models are what most downstream users — structural bioinformaticians, drug discovery teams, textbook and database curators — actually want. Choosing the correct archive for a given data type, rather than depositing everything as PDB supplementary material, is itself a repository-selection decision with real consequences for whether the data is findable and reusable downstream.
Reuse: Structural Bioinformatics, Drug Discovery, and Computed Models
The PDB’s role as reusable infrastructure extends well past the original depositing lab. Its structures underpin structure-based drug design and virtual screening pipelines in pharmaceutical and biotech research, comparative modeling and homology-based structure prediction, structural genomics and protein family classification, and structural biology education. The PDB archive was also the primary training data for the current generation of deep-learning structure predictors, including AlphaFold — a widely cited illustration of how open, well-curated, machine-readable domain repositories compound in value well beyond their original deposit-and-cite purpose.
It’s worth distinguishing predicted from experimental structures explicitly, because the distinction is a frequent source of confusion: the primary PDB archive holds experimentally determined structures. Computed structural models — predictions from tools such as AlphaFold, or other computational modeling methods — are not deposited into the primary PDB archive as if they were experimental results. They are catalogued through separate, purpose-built resources (for example, the AlphaFold Protein Structure Database and community model archives), which wwPDB and partner infrastructure increasingly federate with and link to, while keeping the provenance distinction between “experimentally observed” and “computationally predicted” structures explicit and machine-readable.
What This Means for Data Management Plans and Funder Compliance
For a research administrator reviewing or drafting a DMP that names the PDB as the intended repository for structural output, a few things are worth checking directly rather than assuming: that the anticipated experimental method (X-ray, NMR, cryo-EM) is one the PDB natively accepts and validates; that raw data outputs (images, tilt series) are routed to the correct companion archive (EMPIAR/EMDB) rather than assumed to be covered by a PDB deposition; that any embargo period named in the plan is consistent with wwPDB’s standard hold options; and that the PDB is described accurately as a certified, community-endorsed domain repository rather than a generic file store — which matters when a DMP is being assessed against FAIR data and repository-trustworthiness criteria. Because deposition and access are free of charge and release is effectively permanent once an embargo lapses, the PDB is generally a low-friction, low-risk choice to name in a DMP for standard experimentally determined macromolecular structures — the more common planning error is misrouting raw data that the PDB itself does not accept.
Frequently Asked Questions
Is depositing a structure in the PDB mandatory?
There is no single universal legal mandate, but in practice it functions as one: most journals publishing structural biology results require PDB deposition and public release of an accession code as a condition of publication, and many funders’ data-sharing policies expect structural coordinates to be deposited in a recognized public archive like the PDB.
Is there a cost to deposit or access PDB data?
No. wwPDB states that deposition, validation, biocuration, and remediation services are provided free of charge, and released PDB entries are freely and openly accessible to anyone.
What is the difference between the PDB, EMDB, and EMPIAR?
They archive different stages of the same cryo-EM workflow: EMPIAR holds raw 2D image data, EMDB holds the derived 3D density maps, and the PDB holds the final, interpreted atomic model. See CASRAI’s EMPIAR guide for the full breakdown.
Can AlphaFold-predicted structures be deposited in the PDB?
The primary PDB archive is for experimentally determined structures. Computationally predicted models are catalogued separately (for example, in the AlphaFold Protein Structure Database) rather than deposited into the PDB as though they were experimental entries, though wwPDB and partner infrastructure increasingly link the two.
What file format does the PDB use today?
PDBx/mmCIF is wwPDB’s current master format. The legacy fixed-column “PDB format” is still generated for many entries for backward compatibility, but is a derived output rather than the authoritative source.
Related CASRAI Resources
- Data Repository — the general concept the PDB is an instance of.
- Persistent Identifier (PID)
- FAIR Data Principles
- Data Management Plan (DMP)
- EMPIAR: The Electron Microscopy Public Image Archive
- How to Choose an Open Data Repository
- Research Data Management — Pillar Overview







