Repositories, Preservation & Storage Infrastructure
Once data is ready to be shared or archived, it needs somewhere to live that will keep it accessible and intact over time. This sub-cluster covers the infrastructure side of research data management: how repositories are chosen, evaluated, and certified, and what technical and organizational commitments underlie long-term digital preservation. CoreTrustSeal certifies repositories against sixteen requirements spanning organizational infrastructure, digital object management, and technology, giving funders and researchers a way to judge whether a given repository is trustworthy before depositing data with it. The Digital Curation Centre's (DCC) Curation Lifecycle Model frames preservation and storage as ongoing actions rather than a one-time deposit, with distinct steps for preservation action (migrating formats, fixing bit rot) and store (maintaining the technical environment that keeps data usable). The re3data registry, meanwhile, provides a searchable directory of repositories with structured criteria covering subject coverage, access conditions, and certification status, which researchers and administrators use to identify an appropriate repository for a given dataset or funder mandate. For a research administrator, this area is where DMP commitments become operational reality: a plan that names a repository is only credible if that repository meets certification standards and will still exist, and remain accessible, well past the end of the award. Pages in this sub-cluster cover how to evaluate a repository against these criteria, what preservation actions a certified repository is expected to perform, and how institutional versus disciplinary versus generalist repositories differ in scope and guarantees.
Guides
DOE PAGES: The Department of Energy’s Public Access Repository
DOE PAGES (Public Access Gateway for Energy & Science) is the Department of Energy’s designated repository for peer-reviewed manuscripts and articles from DOE-funded research. This guide covers who must deposit, the E-Link workflow, embargo history, and compliance mechanics for grant and contract administrators.
USGS ScienceBase: Trusted Digital Repository Status, Deposit Eligibility, and DMP Fit
Who can actually deposit in USGS ScienceBase, how its Trusted Digital Repository status is certified against CoreTrustSeal criteria, and when it belongs on a data management plan versus a domain or generalist repository.
Environmental Data Initiative (EDI): The LTER Network’s Data Repository
EDI is the repository that publishes LTER Network and broader environmental/ecological data as EML-described data packages — distinct from the EML standard itself and from NEON’s data-generating role.
Crystallography Open Database (COD): Open-Access Crystal Structure Repository
The Crystallography Open Database (COD) is a free, open-access repository of experimentally determined crystal structures in CIF format, with 520,000+ entries.
Paleobiology Database (PBDB): Fossil Occurrence & Taxonomic Data Repository
The Paleobiology Database (PBDB) is a free, community-curated repository of fossil occurrence and taxonomic data, covering the Phanerozoic under a CC BY 4.0 license and interoperable with GBIF via Darwin Core.
PubChem: NIH/NLM’s Chemistry Compound and Bioassay Data Repository
PubChem is NCBI/NLM/NIH’s free repository of chemical structures, substances, and bioassay data. What it covers, how it’s organized, and how research data managers use it.
CUAHSI HydroShare: The Collaborative Repository for Hydrology Data
HydroShare is CUAHSI’s NSF-funded domain repository for hydrology and water-resources data, models, and time series. Learn what it supports, how DOIs and publication work, and when it’s the right DMP choice.
DARIAH-EU: Europe’s Digital Research Infrastructure for Arts and Humanities
DARIAH-EU is the European Research Infrastructure Consortium (ERIC) supporting digital arts and humanities research. Learn its governance, member countries, funding model, and key services like the SSH Open Marketplace.
The Argo Float Program: Open Oceanographic Data Infrastructure
A guide to the Argo profiling-float program: how ~4,000 autonomous floats collect ocean temperature and salinity data, how it is governed, quality-controlled, and openly shared, and what it illustrates about large-scale open research data infrastructure.
IRIS/EarthScope: The SAGE Seismological Data Archive
How the EarthScope Consortium-operated SAGE Facility (formerly IRIS) archives global seismological and geophysical data, the standards it uses, and what the 2023 IRIS/UNAVCO merger means for citing it.
InvenioRDM: The Open-Source Repository Platform Behind Zenodo
InvenioRDM is the free, open-source repository application that Zenodo itself runs on. This guide covers what it is, who governs it, which institutions run their own instance, and how it compares to DSpace, EPrints, and Fedora.
GenBank: NCBI’s Foundational DNA and Genetic Sequence Database
GenBank is NCBI’s open, annotated archive of all publicly available DNA and RNA sequences — the base nucleotide layer that many other genomics repositories build on or reference. This guide covers what GenBank is, how it differs from RefSeq, how to submit and cite a sequence, and how it relates to the more specialized repositories already covered in this cluster.
HEPData: The Particle-Physics Research Data Repository
HEPData is the discipline-specific repository for scattering and kinematic data underlying particle-physics publications, hosted at Durham University and tightly linked to arXiv and INSPIRE-HEP.
CKAN: Open-Source Data Portal Software for Research and Government Data Catalogs
CKAN is free, open-source data portal software institutions self-host to catalog and publish datasets, used by many government open-data portals and some research data catalogs.
TalkBank and CHILDES: The Domain Repository for Language Acquisition and Communication Research
TalkBank is a research infrastructure for studying human communication; CHILDES, its flagship child-language-acquisition component, is the field-standard repository for language development data. This guide covers governance, the CHAT transcription format, licensing, and how it compares to CLARIN.
CyVerse: Cyberinfrastructure and Data Repository for Plant Biology and Life Sciences
CyVerse is an NSF-funded, open-source cyberinfrastructure platform providing data storage, bioinformatics analysis tools, and DOI-issuing dataset publication for plant biology and life-sciences research.
Recherche Data Gouv: France’s National Research Data Infrastructure
Recherche Data Gouv is France’s national ecosystem for managing, preserving, and sharing research data, combining a central platform with a nationwide competence-centre support network.
DesignSafe-CI: The NHERI Data Repository for Natural Hazards Engineering
DesignSafe-CI is the NSF-funded data repository and cyberinfrastructure for NHERI, hosting experimental, simulation, and field-reconnaissance data for earthquake, wind, tsunami, and storm-surge engineering research. What it covers, how curation and DOI publication work, and where it fits an NSF data management plan.
Ag Data Commons: USDA’s Repository for Agricultural Research Data
What Ag Data Commons is, what data types it hosts, how it fulfills USDA’s public-access requirements, and how it fits among federal domain repositories.
CLARIN: Europe’s Research Infrastructure for Language Resources
CLARIN is the pan-European ERIC infrastructure for language data and processing tools, delivered through a federation of national consortia and built around FAIR-compliant access for linguistics and digital-humanities research.
ImmPort: NIAID’s Data Repository and Analysis Portal for Immunology Research
ImmPort is NIAID/DAIT’s public repository and analysis platform for immunology research data — assay results, clinical study data, and HIPC datasets — and the discipline-specific repository NIH points to for immunology data before a generalist option.
OBIS: The Ocean Biodiversity Information System for Marine Research Data
OBIS is the global open-access platform for marine biodiversity occurrence data, governed under IOC-UNESCO — related to, but organizationally and scientifically distinct from, GBIF.
UniProt: The Protein Sequence and Functional Annotation Knowledgebase
UniProt is the primary knowledgebase for protein sequence and functional annotation data, curated by a consortium of EMBL-EBI, SIB, and PIR. This guide covers UniProtKB’s Swiss-Prot/TrEMBL split, governance, licensing, and how it differs from the Protein Data Bank.
Protein Data Bank (PDB): Archiving and Reusing Macromolecular Structure Data
How the Protein Data Bank archives experimentally determined 3D macromolecular structures, wwPDB governance, deposition mandates, formats, and reuse in structural biology and drug discovery.
CDS Strasbourg: VizieR, SIMBAD, and Aladin as Astronomical Data Services
CDS Strasbourg operates SIMBAD, VizieR, and Aladin — the astronomical object database, catalogue service, and interactive sky atlas that much of professional astronomy relies on for FAIR-compliant data discovery, citation, and reuse. This guide explains what each does and how they interoperate.
NASA Planetary Data System (PDS): Archiving and Reusing Planetary-Mission Science Data
How NASA’s Planetary Data System archives and shares planetary-mission science data: its federated node structure, the PDS4 data standard, data discovery, citation practice, and its fit within FAIR/open-science compliance.
DANS: The Netherlands’ National Research Data Archive and DMP Support
DANS is the Netherlands’ national institute for research data archiving, jointly run by KNAW and NWO. This guide covers its Data Stations, DataverseNL, CoreTrustSeal certification, and how it fits into Dutch DMP compliance.
PANGAEA: Data Publisher for Earth and Environmental Science Datasets
PANGAEA is an open-access data publisher for earth, environmental, polar, and marine science data, offering DOI-based citation, editorial curation, and journal-linked data papers.
Earth System Grid Federation (ESGF): Federated Access to CMIP Climate Model Data
What the Earth System Grid Federation (ESGF) is, how it distributes CMIP climate model data underlying IPCC assessment reports, and how researchers search, download, and manage access to it.
Copernicus Climate Data Store (C3S): Accessing Europe’s Climate Reanalysis and Projection Data
How to access the Copernicus Climate Data Store (CDS): what the Copernicus Climate Change Service (C3S), operated by ECMWF, provides — ERA5 reanalysis, regional reanalyses, and CMIP/CORDEX climate projections — plus the registration, API, and licensing steps researchers and RDM offices need.
GWOSC: Public Access to LIGO, Virgo, and KAGRA Gravitational-Wave Data
GWOSC is the public repository for LIGO, Virgo, and KAGRA gravitational-wave strain data and event catalogs, including the O4b release and GWTC-5.0 catalog.
NOAA NCEI: Accessing the National Centers for Environmental Information’s Climate and Ocean Data
How to access NOAA’s National Centers for Environmental Information (NCEI): its role as NOAA’s climate, weather, and ocean data archive, its predecessor agencies, what it holds, and the access tools researchers actually use.
NEON (National Ecological Observatory Network): Continental-Scale Ecological Data for Researchers
NEON is an NSF-funded, Battelle-operated observatory producing standardized, long-term ecological and environmental sensor data across 81 U.S. field sites — distinct from occurrence-record aggregators like GBIF. This guide covers what NEON measures, how it differs from GBIF, and how to access and cite its data.
NASA Exoplanet Archive: The Domain Repository for Confirmed and Candidate Exoplanet Data
A guide to the NASA Exoplanet Archive: what it is, who operates it, its confirmed/candidate planet and host-star data, its mission tables and tools, and how to cite its continuously updated datasets.
Materials Genome Initiative (MGI): The US Push for FAIR Materials-Science Data
The Materials Genome Initiative (MGI) is a US federal multi-agency effort, launched in 2011 and coordinated via NIST and the NSTC, to accelerate materials discovery through integrated computation, experiment, and FAIR data infrastructure.
MetaboLights: EMBL-EBI’s Open-Access Metabolomics Data Repository
What MetaboLights is, who operates and funds it (EMBL-EBI/ELIXIR), what data types and ISA-Tab metadata it requires, how MTBLS accessions work, and how it relates to Metabolomics Workbench.
EMPIAR: The Electron Microscopy Public Image Archive for Raw Cryo-EM Data
EMPIAR is EMBL-EBI’s public archive for raw 2D cryo-EM image data. Learn how it differs from EMDB (derived maps) and the PDB (atomic models), and how to submit.
ArrayExpress: EBI’s Functional-Genomics/Transcriptomics Data Archive
ArrayExpress is EMBL-EBI’s archive for functional-genomics data, now operating as part of BioStudies. This guide covers what it hosts, the 2021 BioStudies migration, submission via Annotare, accession-number conventions (E-MTAB, E-GEOD, E-MEXP), and its relationship to NCBI’s GEO.
Metabolomics Workbench: The NIH-Funded National Metabolomics Data Repository
What Metabolomics Workbench is, who funds and operates it, what data it accepts, its mwTab submission requirements, its Study/Analysis ID system, and how it relates to MetaboLights.
Gene Expression Omnibus (GEO): NCBI’s Repository, MIAME/MINSEQE, and Accession Numbers
GEO is NCBI’s public repository for gene-expression data, built around the MIAME/MINSEQE minimum-information standards and a GSE/GSM/GPL/GDS accession system.
WorldPop: Open Gridded Population Data for Research
What WorldPop is, the public health, epidemiology, and disaster-response research it supports, its GeoTIFF gridded data formats and open access model, CC BY 4.0 licensing, and how to correctly cite a WorldPop dataset.
OpenNeuro: The BIDS-Based Neuroimaging Data-Sharing Platform
What OpenNeuro is, how it enforces the BIDS standard, which neuroimaging data types it hosts, and how it fits into the broader neuroimaging data-sharing ecosystem.
GREI: The NIH Generalist Repository Ecosystem Initiative
GREI is NIH ODSS’s 2022 initiative coordinating seven generalist repositories (Dataverse, Dryad, Figshare, Mendeley Data, OSF, Vivli, Zenodo) around shared standards for NIH-funded data sharing.
CESSDA: The Consortium of European Social Science Data Archives Explained
CESSDA is the European Research Infrastructure Consortium (ERIC) that federates national social-science data archives across 22 member countries behind a shared catalogue, trust framework, and FAIR-aligned tools.
Data Rescue Projects: Recovering At-Risk Scientific Data Before It’s Lost
What data rescue projects are, why at-risk research and environmental data gets lost, and how real initiatives like Data Refuge, DataRescue, and IEDRO have recovered it.
The “Desirable Characteristics of Data Repositories” Checklist Explained
A section-by-section walkthrough of the NSTC’s 2022 federal guidance document: all 14 desirable characteristics for data repositories, plus the 7 additional considerations for repositories storing human data, and how research administrators actually use it.
SSRN (Social Science Research Network): How It Works for Researchers
A guide to how SSRN (Social Science Research Network) actually works: what researchers can post, how its subject-matter research networks and eJournals are organized, the submission and review process, its ownership by Elsevier, and how an SSRN posting differs from a peer-reviewed journal article.
Research Repositories: What Counts as One, and How They Differ From Archives
Not every place research output is stored qualifies as a repository. Here is what distinguishes a research repository from a general archive, institutional repository, or preprint server — and how to choose between them.
Data Publication Practices: Repositories, Data Journals, and Data Papers
What formally “publishing” a research dataset requires beyond sharing or depositing it: choosing between a repository DOI and a data-journal data paper, metadata, licensing, versioning, and how dataset peer review differs from manuscript review.
arXiv Preprints: What They Are and How to Use Them
A practical guide to arXiv: subject categories, the endorsement and moderation system, how it differs from peer-reviewed publication, and how to cite an arXiv preprint correctly.
NSF Public Access Repository (NSF-PAR): Requirements, Deposit Process & Public Access Plan 2.0
NSF-PAR is the National Science Foundation’s repository for peer-reviewed publications from NSF-funded research. This guide covers deposit requirements, the 2015 and 2023 Public Access Plans, the 2026 move to zero embargo, and how NSF-PAR differs from NIH’s PubMed Central mechanism.
ISBER vs. NCI Best Practices for Biorepositories: What’s the Difference?
A comparison of the two leading biorepository best-practices documents: ISBER’s Best Practices for Repositories and NCI’s Best Practices for Biospecimen Resources, and how institutions use them together.
Globus Data Transfer: How It Works and Why Research Institutions Use It
Globus is a nonprofit research data transfer service, built on the GridFTP protocol and operated by the University of Chicago, that lets researchers move large datasets reliably between institutions, HPC systems, and repositories with automatic retry and checksum verification.
Genome Projects and Research Data Management
How the Human Genome Project, 1000 Genomes, and GenBank shaped modern research data management, from the Bermuda Principles to today’s open/controlled-access repository model.
Dryad Data Repository: Costs, Curation, and DMP Fit
A practical look at Dryad’s curated submission workflow, its CC0-default licensing, current Data Publishing Charges, and where it fits relative to Zenodo, Figshare, and institutional repositories when writing a data management plan.
Figshare Data Repository: Features, Certification, and DMP Fit
A practical look at figshare as a generalist data repository: features, certification status (ISO 27001 vs CoreTrustSeal), and how to use it in a DMP.
CoreTrustSeal Certification: Requirements, Process, and Whether It’s Worth Pursuing
The full 16 CoreTrustSeal requirements, the self-assessment and peer-review process end to end, real certified-repository examples, and practical guidance for deciding whether your institution should pursue certification.
Checksum Verification and Fixity Checking for Archived Research Data
How checksum verification and fixity checking work, why trustworthy-repository frameworks like CoreTrustSeal and the NDSA Levels of Digital Preservation require them, and the tools repositories use to detect silent corruption in archived research data.
How to Choose an Open Data Repository
Choosing an open data repository is a DMP commitment that also determines whether your deposit is discoverable and reportable through your institution’s CRIS.







