Examples
Worked examples
- Is an instance
A university natural history collection maps its specimen catalog to Darwin Core terms and packages it as a Darwin Core Archive using GBIF's Integrated Publishing Toolkit (IPT). Once registered with its national GBIF Participant Node, the records become searchable through GBIF.org, receive a GBIF dataset DOI, and are indexed alongside occurrence records from thousands of other institutions.
- Is an instance
A researcher runs an occurrence search on GBIF.org filtered to a genus and a country, then downloads the result set. GBIF issues a citable 'download DOI' that snapshots the exact query and the version of every constituent dataset at that moment, which the researcher cites in their methods section -- this is the standard way GBIF-mediated data gets cited in a publication, distinct from citing any single contributing dataset directly.
Counter-examples
Looks similar, but isn't
- Not an instance
A national biodiversity agency maintains its own internal occurrence database, structured using Darwin Core terms, but has never registered a dataset with a GBIF Participant Node or published it via the IPT/DwC-A pathway. This is not GBIF data, even though it uses the Darwin Core standard -- Darwin Core is the data standard; GBIF is one (the largest) network that aggregates data described in that standard, and adopting the standard does not by itself make a dataset part of the GBIF-mediated network.
- Not an instance
A marine-biodiversity dataset published only through OBIS (the Ocean Biodiversity Information System), a separate Darwin Core-based network with its own node structure and registration process, is not GBIF-mediated unless it is separately registered with a GBIF node -- OBIS and GBIF are parallel, interoperable aggregators built on the same underlying standard, not the same platform.
Editorial commentary
GBIF — the Global Biodiversity Information Facility — is an intergovernmental network and open-access data infrastructure that aggregates species-occurrence and specimen records from data-holding institutions worldwide and makes them freely searchable, downloadable, and citable through a single interface at GBIF.org. It does not define its own data standard; instead it is built on top of Darwin Core, the biodiversity-data standard maintained by Biodiversity Information Standards (TDWG), and is best understood as the largest platform that consumes and republishes Darwin Core-formatted records, not as a standard in its own right.
How data reaches GBIF
Data does not appear in GBIF automatically. An institution — a natural history museum, herbarium, university collection, government monitoring program, or citizen-science platform — maps its records onto Darwin Core terms and packages them as a Darwin Core Archive (DwC-A): a data file (or files) plus a machine-readable meta.xml descriptor mapping each column to its Darwin Core term, typically alongside Ecological Metadata Language (EML) dataset-level metadata. The publisher most commonly builds this archive using GBIF’s own Integrated Publishing Toolkit (IPT), a free, widely deployed open-source tool, though other DwC-A-compliant publishing software can also register with the network. Once the dataset is registered through a national or thematic GBIF Participant Node, GBIF’s indexing pipeline harvests, validates, and periodically re-crawls it, and the records become part of the searchable occurrence index.
Governance: the Participant Node network
GBIF is an intergovernmental initiative, not a single company or repository operator. It is coordinated by a Secretariat based in Copenhagen and funded collectively by its participating governments and organizations, but the actual data holdings are contributed and managed through a distributed network of national and thematic Participant Nodes — designated teams that coordinate the institutions producing and publishing biodiversity data within their own country or thematic community. Since 2008 this network has been organized into regional groupings spanning Africa, Asia, Europe and Central Asia, Latin America and the Caribbean, North America, and Oceania. This node structure is why GBIF functions as an aggregating platform rather than a single centralized data provider — the underlying records remain the responsibility of, and are cited back to, the original publishing institution.
Occurrence search, downloads, and citation
The most common way a researcher interacts with GBIF is through its occurrence-search interface or API: filtering by taxon, geography, time period, basis of record, or dataset, then requesting a download of the matching records. GBIF issues each such download a persistent, citable ‘download DOI’ that snapshots the exact query and the version of every constituent dataset at that moment — the standard mechanism for citing GBIF-mediated data in a publication’s methods section, and one that supports reproducibility since the snapshot does not change even as the live index continues to grow.
GBIF, Darwin Core, and OBIS — how the pieces relate
These three terms are frequently conflated but describe different layers. Darwin Core is the data standard — the controlled set of terms (scientificName, decimalLatitude, eventDate, and the rest) that a biodiversity dataset’s fields are mapped onto. GBIF is a platform that aggregates Darwin Core-formatted datasets from a global network of contributing institutions and makes them jointly searchable. The Ocean Biodiversity Information System (OBIS) is a comparable, parallel platform — also built on Darwin Core, but focused on marine biodiversity data and operating its own separate node and registration structure. A dataset can, and often does, use the Darwin Core standard without ever being registered with either network; using the standard is a precondition for aggregation, not the same thing as being aggregated. See CASRAI’s guide to choosing a metadata schema for a dataset for the general principle of matching a schema to the repository or aggregator that will actually consume it, and the DataCite metadata schema entry for the separate, general-purpose schema most repositories use for the DOI-registration metadata that sits alongside — not instead of — the Darwin Core-formatted occurrence data itself.
References
- GBIF Secretariat, ‘What is GBIF?’ (gbif.org/what-is-gbif).
- GBIF, ‘GBIF nodes’ and ‘Establishing an Effective GBIF Participant Node’ documentation (gbif.org, docs.gbif.org).
- GBIF Integrated Publishing Toolkit (IPT) user manual, ‘Darwin Core Archives — How-to Guide’ (ipt.gbif.org/manual/en/ipt/latest/dwca-guide).
- TDWG (Biodiversity Information Standards), Darwin Core normative term list (dwc.tdwg.org).
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="GBIF (Global Biodiversity Information Facility)"
vocab-term-identifier="https://casrai.org/dictionary/term/gbif-global-biodiversity-information-facility" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/gbif-global-biodiversity-information-facility",
"name": "GBIF (Global Biodiversity Information Facility)",
"identifier": "https://casrai.org/dictionary/term/gbif-global-biodiversity-information-facility",
"description": "GBIF (the Global Biodiversity Information Facility) is an intergovernmental network and open-access data infrastructure -- coordinated by a Secretariat in Copenhagen and funded by its participating governments and organizations -- that aggregates, indexes, and republishes species-occurrence and specimen records contributed by data-holding institutions worldwide through a global network of national and thematic Participant Nodes. A dataset counts as GBIF-mediated when it has been registered through a GBIF Participant Node (or directly via GBIF.org) and is discoverable, searchable, and downloadable through the GBIF occurrence-search interface and API; the great majority of datasets reach GBIF as Darwin Core Archives (DwC-A), most commonly built and published using GBIF's own Integrated Publishing Toolkit (IPT). GBIF itself does not define a data standard -- it consumes Darwin Core, the pre-existing TDWG biodiversity-data standard -- and is best understood as the largest aggregating platform built on top of that standard, not the standard itself.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
"url": "https://casrai.org/dictionary/term/gbif-global-biodiversity-information-facility",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"dateModified": "2026-07-24T06:17:35",
"inLanguage": "en"
}






