Most guidance on choosing an open data repository stops at a checklist: check your funder’s policy, prefer a discipline-specific repository where one exists, fall back to a generalist repository, look for a DOI, check the storage limit and the licence. That checklist is correct as far as it goes, and CASRAI’s own Research data infrastructure domain defines the vocabulary behind it — generalist repository, discipline-specific repository, trusted digital repository, and 40 other terms. This guide does not repeat that definitional work. It covers the two things a generic research-services checklist won’t: how repository choice functions as a specific, trackable commitment inside a data management plan (DMP), and how a repository’s metadata determines whether the deposit ever becomes visible to your institution’s research information system (CRIS) — CASRAI’s actual area of standards work.
Repository choice is a DMP commitment, not an isolated decision
Picking a repository is rarely a standalone task. In a funder- or institution-mandated Data Management Plan, the named repository is usually recorded against a specific DMP component — most often the sharing commitment and, separately, the preservation commitment, since a repository chosen for discoverability during a project is not always the one responsible for long-term retention after it ends. Naming a repository in a DMP is a forward commitment that a compliance check, a funder audit, or your institution’s data steward can later verify against — so the repository named at proposal stage should be one you can actually deposit into under the terms described.
This is also where DMPs stop being static PDFs. Modern DMP tools — DMPonline (used for UKRI, Horizon Europe, and Wellcome templates) and Argos (OpenAIRE’s tool) among them — increasingly produce a machine-actionable DMP (maDMP), following the RDA DMP Common Standard. In a maDMP, the repository isn’t just a name in a text field — it’s referenced by its re3data identifier, alongside persistent identifiers for the project, the people (ORCID), and the funder. That’s what lets a system auto-populate a repository deposit form, or check “planned repository” against “actual deposit location” at compliance-reporting time. A repository that isn’t re3data-listed, or doesn’t expose a stable identifier, is harder to reference this way — worth knowing before it’s the one named in a live DMP.
Discipline-specific vs. generalist: the first fork, briefly
CASRAI’s dictionary already distinguishes discipline-specific repositories (GenBank/EMBL/DDBJ for nucleotide sequences, PDB for protein structures, PANGAEA for earth-science observation data, and similar domain-curated destinations) from generalist repositories (Zenodo, Figshare, Dryad, Harvard Dataverse, OSF), which accept any output type and mint a DataCite DOI on deposit. Funders that maintain approved-repository lists (NIH’s is a commonly cited example) generally expect a discipline-appropriate repository first, and only fall back to a generalist one when no suitable domain repository exists. If your data involves controlled or human-subjects access rather than open download, that’s a different category again — a sensitive-data repository such as the European Genome-phenome Archive or dbGaP, operating under a Trusted Research Environment model rather than open deposit.
To find the specific repository rather than the category, use the registries built for exactly that: re3data.org (the Registry of Research Data Repositories, indexing 3,000+ repositories with structured metadata on discipline, access conditions, certification, and API availability) and FAIRsharing.org, which cross-links repositories to the metadata standards and data policies that reference them. CASRAI’s dictionary carries both as terms in their own right — re3data and FAIRsharing — because funders and journals routinely write deposit policy in terms of “a re3data-listed repository,” making registry presence itself a practical eligibility signal, not just a convenience for searching.
Verifying trust: what “certified” actually means
“Trusted repository” is a specific, third-party-assessed status, not a marketing claim. Three certification tiers exist at different depths: CoreTrustSeal — a lightweight, peer-reviewed self-assessment against 16 published requirements, renewable every three years, and by far the most widely held baseline (formed in 2017 from the merger of the Data Seal of Approval and the ICSU World Data System certification, per CASRAI’s CoreTrustSeal and World Data System certification entries); the German nestor seal (DIN 31644); and ISO 16363, an external-audit standard that is more rigorous, and rarer, than either. An increasing number of funders and the European Commission’s EOSC and OpenAIRE-Nexus programmes require or strongly prefer deposit into a CoreTrustSeal-certified (or equivalent) repository — worth checking directly against the current certificate list rather than assuming a repository’s reputation implies certification.
Metadata schema: the difference between “deposited” and “discoverable”
This is the step a generic repository checklist usually skips, and it’s where CASRAI’s standards work is most directly relevant. Depositing a dataset does not, by itself, make it discoverable outside the repository’s own search box. What does is the metadata record the repository generates on deposit, and the protocol it uses to expose that record for harvesting.
Most generalist repositories, and a growing share of discipline-specific ones, generate records against the DataCite Metadata Schema (kernel 4; version 4.7 is current, though plenty of production repositories still validate against 4.4 or 4.5) when they mint a DOI. That schema’s mandatory and optional properties — Creator, Title, PublicationYear, ResourceType, FundingReference, RelatedIdentifier, and others — are what get exposed for harvesting, typically over OAI-PMH, the twenty-year-old but still-dominant harvesting protocol behind OpenAIRE, BASE, CORE, and most national aggregators. A dataset with a thin metadata record (no FundingReference, no RelatedIdentifier back to the publication it supports) is technically deposited but effectively invisible to anything that harvests on those fields — including your own institution’s reporting systems, covered next.
Two related terms worth knowing before you evaluate a repository’s metadata output: a dataset landing page (the citable, human-readable record a DOI resolves to) and a data publication platform (repositories or data journals — PANGAEA, Dryad’s journal partnerships, Scientific Data — that treat the dataset itself as a citable, peer-reviewable output rather than a passive file store).
From deposit to your institution’s CRIS
This is the part that rarely appears outside CASRAI’s own domain, and it’s the second half of why repository choice matters beyond the deposit itself. Most research-intensive institutions run a Current Research Information System (CRIS) — DSpace-CRIS, Pure, or a similar platform — that reports outputs, funded projects, and researcher activity up to funders, national assessment exercises, and institutional dashboards. A dataset deposited in an open repository does not automatically appear there. It has to be linked to a project record, a researcher’s affiliation at the time of the work, and often a research activity record.
The mechanism that makes this linking automatic rather than manual is CRIS interoperability: shared models — chiefly CERIF, maintained by euroCRIS — shared exchange protocols (OAI-PMH, REST), and shared identifiers (ORCID, ROR, DOI, RAiD) that let a dataset record minted by a repository be resolved unambiguously against a person, an organisational unit, and a project already sitting in the CRIS. Platforms like DSpace-CRIS implement this directly, treating people, projects, and organisations as first-class entities alongside deposited items. The OpenAIRE Guidelines — one profile for data archive managers (built on the DataCite schema) and a separate one for CRIS managers (profiling CERIF) — exist specifically to keep these two sides compatible.
The practical consequence for repository choice: a repository that exposes rich, identifier-carrying metadata over a standard harvesting protocol lets your institution’s CRIS pull the deposit in with little or no manual re-entry. A repository that doesn’t creates a second data-entry job for whoever maintains your institutional research record — every time, for every deposit.
A decision checklist tied to what actually differs
- Named in your DMP’s sharing or preservation commitment? If the DMP is or will become machine-actionable, is the repository identifiable by a re3data ID your DMP tool (DMPonline, Argos) can resolve?
- Discipline-specific option first. Check re3data.org or FAIRsharing.org by subject before defaulting to a generalist repository.
- Certification. CoreTrustSeal (or ISO 16363/nestor seal) status, checked against the current certificate list, not assumed from reputation.
- Metadata output. Does it mint a DataCite DOI with FundingReference and RelatedIdentifier populated? Does it expose records over OAI-PMH or a documented API?
- CRIS path. Can your institution’s CRIS harvest from this repository automatically (via CERIF-XML, OAI-PMH, or a documented connector), or will the deposit need manual re-entry into your institutional research record?
Frequently asked questions
What are some examples of open data repositories?
Generalist: Zenodo, Figshare, Dryad, Harvard Dataverse, OSF. Discipline-specific: GenBank/EMBL/DDBJ (nucleotide sequences), PDB (protein structures), PANGAEA (earth-science observation data), and the UK Data Service (social and economic survey data). Which one is appropriate depends on your discipline and your funder’s or journal’s deposit policy — see the discipline-specific vs. generalist section above.
What’s the difference between a “scientific data repository” and a generalist one?
A discipline-specific (sometimes called “domain” or “scientific”) repository applies field-appropriate curation, standardised metadata, and controlled vocabularies specific to that discipline — GenBank’s sequence-format requirements, for example. A generalist repository accepts any file type with minimal domain curation, trading depth for breadth. CASRAI’s dictionary treats “domain repository” and “discipline-specific repository” as largely interchangeable terms.
What is a “registry of research data repositories,” and is it a repository itself?
No — re3data.org and FAIRsharing.org are directories, not places you deposit data. They index thousands of actual repositories with structured metadata (discipline, access conditions, certification, PID systems, API availability) so you can search for and compare the repositories that fit your data, then deposit into the repository you select, not into the registry.
Does depositing in an open data repository automatically satisfy my funder’s DMP requirement?
Not by itself. Most DMP mandates require the repository to be named as a commitment and to meet specific conditions — certification status, embargo support, an acceptable licence — that the deposit then has to fulfil in practice. A mismatch between what the DMP promised and where the data actually ended up is exactly what a funder compliance check looks for.
Does an open data repository deposit automatically show up in my institution’s CRIS?
Only if the repository exposes metadata your CRIS can harvest and the record carries identifiers (ORCID, project ID, funder ID) your CRIS can resolve. Otherwise, someone has to re-enter the deposit into the institutional research record by hand. See the CRIS section above for what makes this automatic.







