The Paleobiology Database (PBDB) is a free, public, community-curated repository of fossil occurrence and taxonomic data. Instead of holding raw specimen images or files, PBDB stores structured records: where and when a fossil taxon was found, how it was identified, and how that identification has changed over time. For research administrators and data stewards working with paleontology, geology, or evolutionary-biology projects, PBDB is the field’s closest equivalent to a domain repository — a discipline-specific alternative to depositing fossil occurrence data in a generalist repository such as Zenodo or Figshare.
What PBDB actually stores
PBDB’s core unit is the occurrence: a record that a particular fossil taxon was found at a particular collection locality, within a defined stratigraphic and temporal interval. Around that core, the database holds several linked record types:
- Fossil collections — the locality-level metadata (geographic coordinates, stratigraphic unit, geologic time interval, lithology, collection method) that occurrence records are attached to.
- Taxonomic names and opinions — the nomenclatural history of each taxon, including synonymies and reclassifications, so a search can resolve outdated names to their current usage.
- Taxonomic re-identifications — a record of how a given occurrence’s identification has changed as expert opinion or evidence changed, rather than a single fixed label.
- Bibliographic references — the published source each occurrence and taxonomic opinion is drawn from, so every record is traceable to primary literature.
- Geologic time scales — the chronostratigraphic framework (eras, periods, stages) used to bound occurrences in time.
All of this is held in a relational database and covers animals, plants, and microorganisms across the Phanerozoic (roughly the last 540 million years). Because occurrences are geographically and temporally explicit, the aggregate dataset supports macroevolutionary and paleoecological research — diversity curves through time, range shifts, extinction-selectivity analysis — that a single fossil collection or paper could not support on its own.
History and governance
PBDB grew out of the NCEAS (National Center for Ecological Analysis and Synthesis)-funded Phanerozoic Marine Paleofaunal Database initiative, which ran from August 1998 to August 2000, and it received direct National Science Foundation funding from 2000 through 2015. It is not run by a single institution’s library or IT department in the way many domain repositories are; instead it is overseen by an international committee of major data contributors — working paleontologists from institutions worldwide who both enter data and set editorial and access policy. The database’s administrative home has moved between institutions over its history; understanding this contributor-governed structure matters for a data steward evaluating PBDB, since it means data quality and coverage in any given taxonomic group tracks which specialists are actively contributing to that group, not a uniform in-house curation team.
License and how to access the data
PBDB data is released under a CC BY 4.0 license, meaning it can be reused, including commercially, with attribution. There are three main ways to get data out of PBDB:
- PBDB Navigator — a map-based web interface for browsing occurrences interactively by taxon, time interval, and geography.
- Direct downloads — CSV/TSV exports of occurrence, collection, and taxonomic data filtered by search criteria, generated through the main web application.
- The PBDB API — a documented, versioned data service that lets external systems query occurrences, taxa, collections, and references programmatically. The API is described in a peer-reviewed paper (Uhen et al., “The Paleobiology Database application programming interface,” Paleobiology, 2018) and is the mechanism most cyberinfrastructure interoperability with PBDB is built on.
Interoperability: Darwin Core, GBIF, and neighboring databases
PBDB does not operate in isolation from the broader biodiversity-informatics stack. Its dataset is registered and discoverable through the Global Biodiversity Information Facility (GBIF), the main international aggregator for species-occurrence data, which normalizes contributed records into Darwin Core — the Biodiversity Information Standards (TDWG) term set most biodiversity data-sharing infrastructure uses to describe taxon occurrences. That mapping is what lets a fossil occurrence recorded in PBDB’s own paleontology-specific schema surface in GBIF-mediated searches alongside modern-species occurrence data from herbaria, museums, and citizen-science platforms.
PBDB is also part of a small cluster of paleontological and paleoecological data infrastructure that has deliberately built interoperable connections rather than staying siloed:
- Neotoma Paleoecology Database — a sister database focused on Quaternary paleoecological data (pollen, plant macrofossils, and related proxies), with which PBDB helped establish the Earth-Life Consortium, a non-profit umbrella organization supporting free sharing of paleoecological and paleobiological data.
- Macrostrat and iDigBio — NSF-funded projects such as ePANDDA have built interoperable APIs connecting PBDB to these platforms, so a query can move between fossil occurrence data, stratigraphic-column data, and digitized natural-history-collection specimen records without manual reconciliation.
This matters for anyone comparing PBDB to a domain database like OBIS in a different subfield (ocean biodiversity rather than paleontology): the pattern of a discipline-specific database that maintains a controlled schema internally while exposing Darwin Core-mapped records to a generalist aggregator is common across biodiversity informatics, not unique to fossils.
Using PBDB in a data management plan
For a paleontology, sedimentary geology, or evolutionary-biology project with a funder-mandated data management plan, naming PBDB as the target repository for fossil occurrence and taxonomic data is a defensible, field-standard choice — provided the data actually fits PBDB’s scope (identified fossil occurrences with locality and stratigraphic context, not raw imagery, CT scans, or unprocessed specimen data, which typically belong in a separate sample or generalist repository alongside a PBDB deposit). When evaluating or writing that DMP commitment, plan authors should be explicit about:
- Persistent identification — PBDB assigns internal record identifiers, but a DMP should also state whether related outputs (the dataset extract itself, a data paper) will carry a separate DOI, since PBDB’s own record IDs are not DOIs.
- Licensing — CC BY 4.0 is PBDB’s blanket license; a DMP doesn’t need to negotiate this, but should note it so downstream reuse terms are clear to reviewers.
- Curation responsibility — because PBDB relies on a distributed community of specialist contributors rather than in-house curators for every taxonomic group, a DMP citing PBDB should identify who on the project team will enter and vouch for the data’s taxonomic identifications, not assume PBDB itself performs independent verification.
See How to Choose an Open Data Repository for the general criteria (subject fit, certification, persistence guarantees) that apply when deciding between a domain repository like PBDB and a generalist alternative, and How to Make Your Dataset FAIR for how repository choice interacts with the FAIR principles more broadly.
Data quality considerations
As with any large, community-curated discipline-specific repository, data quality and taxonomic currency in PBDB vary by contributor and by taxonomic group, since entries reflect the identifications and stratigraphic ranges that individual specialists have submitted and vouched for rather than a single centralized review process applied uniformly across the whole database. Researchers reusing PBDB data for quantitative macroevolutionary analysis should check the bibliographic reference and contributor history attached to the relevant taxonomic group rather than treating all records as equally current, and should consult PBDB’s own documentation on data completeness metrics where analysis depends on it.
Frequently asked questions
Is the Paleobiology Database the same as GBIF?
No. PBDB is a discipline-specific database for fossil occurrence and taxonomic data, maintained by paleontologists under its own governance and schema. GBIF is a generalist international aggregator that indexes occurrence data, including modern species, from thousands of contributing datasets, PBDB among them, normalized into the shared Darwin Core standard. A record can exist in both: entered and curated in PBDB, then surfaced through GBIF’s aggregated search.
Can I cite a PBDB extract as a dataset?
Yes, in the same way researchers cite extracts from other domain repositories: cite the specific query or download (with its parameters and access date) and, where a DMP or journal requires it, describe the extract’s provenance back to PBDB’s underlying occurrence and reference records, since raw PBDB occurrence records don’t individually carry DataCite-style DOIs.
Does PBDB store fossil images or specimen scans?
No. PBDB’s scope is structured occurrence, taxonomic, and collection-locality data, not media files. Projects generating fossil imagery, CT scans, or 3D models typically deposit those separately in a generalist or imaging-specific repository and link back to the relevant PBDB occurrence or collection record.
Who can contribute data to PBDB?
Data entry is performed by an international community of contributing researchers under PBDB’s own editorial policies, rather than being open to unrestricted public submission; the database is overseen by a committee of major data contributors who set access and entry standards.







