The Generalist Repository Ecosystem Initiative (GREI) is a program the National Institutes of Health’s Office of Data Science Strategy (NIH ODSS) launched in 2022 to bring seven established, discipline-agnostic data repositories into closer alignment with NIH’s data-sharing goals. It matters to anyone writing an NIH Data Management and Sharing (DMS) Plan for one specific reason: when a study’s data doesn’t fit an NIH-designated or discipline-specific repository, GREI’s member repositories are the generalist options NIH has actually invested in making more consistent, better documented, and easier to evaluate — without NIH ever designating a single one of them as required or preferred.
What GREI is, and why NIH created it
GREI launched in September 2022 under NIH ODSS, the office responsible for coordinating NIH’s data science strategy across institutes. Its stated mission has two parts: a primary goal of establishing a common set of cohesive, consistent capabilities, services, metrics, and social infrastructure across participating generalist repositories, and a secondary goal of raising researcher awareness of FAIR (Findable, Accessible, Interoperable, Reusable) data principles and helping investigators put them into practice. The initiative sits downstream of a real gap NIH’s own 2023 Data Management and Sharing Policy exposed: that policy requires a plan for depositing and sharing scientific data, but a large share of NIH-funded data has no natural home in an existing NIH-designated or discipline-specific repository, and would otherwise land in a patchwork of generalist repositories with inconsistent metadata, licensing, and long-term guarantees. GREI funds and coordinates work across the seven repositories to close that consistency gap, rather than building a new NIH-run repository from scratch.
The seven GREI member repositories
GREI comprises seven independently operated, established generalist repositories, each already accepting deposits before GREI existed:
- Dataverse — open-source repository software (Harvard’s own installation, plus many institutional instances) built around structured, citable dataset deposits with versioning.
- Dryad — a curated, nonprofit repository historically tied to journal data-availability policies, with a curation step before publication.
- Figshare — a commercial repository (Digital Science) widely used for figures, datasets, and other research outputs, with institutional-branded instances at many universities.
- Mendeley Data — Elsevier’s data repository, integrated with Mendeley reference management and Elsevier’s journal submission workflows.
- Open Science Framework (OSF) — the Center for Open Science’s free platform combining project management, preregistration, and data/materials sharing.
- Vivli — the only GREI member built specifically for controlled-access clinical trial data rather than open deposit; see CASRAI’s guide to how Vivli’s data request and access-review process works for how it differs operationally from the other six.
- Zenodo — CERN-operated, free, discipline-agnostic repository with strong integration into software-citation and GitHub-archiving workflows.
These repositories differ meaningfully in cost structure, curation model, and access controls — GREI membership signals participation in the initiative’s shared standards work, not that the seven repositories are interchangeable for a given dataset.
What GREI actually does
GREI’s work is coordination and standards-building among the seven member repositories, not data hosting itself. Publicly documented workstreams include developing shared metadata practices so a dataset’s core descriptive fields look consistent regardless of which member repository it’s deposited in, defining common usage metrics so funders and institutions can compare data reuse across repositories on the same terms, producing joint use cases and guidance for researchers on where and how to share NIH-funded data, and running training and outreach — public webinars and workshops — aimed at helping investigators and research administrators understand FAIR data practices and repository selection. NIH ODSS and the member repositories have published GREI fact sheets and workshop summary reports documenting this work as it has progressed since 2022.
Where GREI fits in NIH DMS Policy compliance
NIH’s own guidance on selecting a data repository sets out a clear order of preference, and it is important to be precise about where generalist repositories — and GREI specifically — sit in it:
- NIH-designated data repositories for the specific data type, where one exists (for example GenBank or dbGaP for certain genomic and genotype-phenotype data).
- Other established discipline-specific or data-type-specific repositories, evaluated against NIH’s stated “desirable characteristics” — assignment of a citable, unique persistent identifier (a DOI or accession number) resolving to a stable landing page, a credible long-term sustainability plan, and other FAIR-aligned criteria.
- Generalist or institutional repositories, only when no suitable discipline-specific option exists.
NIH is explicit that it does not recommend or require any single generalist repository, GREI membership included. Being one of the seven GREI repositories is not an NIH designation, a compliance shortcut, or a certification — it means that repository has committed to GREI’s shared-standards work and NIH has funded it to build capacity, not that NIH has approved it for your specific dataset. A DMS Plan that names a GREI-member repository still needs to justify that choice against the same desirable-characteristics criteria NIH applies to any other repository named in a plan, and reviewers and Institute program staff will read the plan the same way regardless of GREI membership.
In practice, GREI membership is still a genuinely useful signal for a research administrator triaging repository options at the third tier of that hierarchy: it means the repository has an active NIH-funded relationship, is working toward consistent metadata and FAIR practices with the other six, and has a track record specifically oriented toward NIH-funded data. That is a meaningfully different starting point than an arbitrary generalist repository with no comparable track record — it just isn’t a substitute for confirming the repository actually meets NIH’s stated criteria for the dataset in question.
Practical considerations when naming a GREI repository in a DMS Plan
- Discipline fit first. Confirm no NIH-designated or discipline-specific repository already exists for the data type before defaulting to a generalist option — see NIH’s own repository-selection guidance above.
- Cost. Deposit costs vary widely across the seven — some (OSF, Zenodo) are free at standard deposit sizes; others (Figshare, Dataverse instances, Mendeley Data) may charge for large deposits or offer institutional-subsidized accounts; Dryad charges a per-dataset publication fee. Confirm current pricing directly with the repository before writing it into a budget or plan.
- Access model. Six of the seven are built around open or embargoed deposit; Vivli is structurally different, built for controlled, request-and-review access to clinical trial individual participant data. If a study needs controlled rather than open access, Vivli (or a comparable controlled-access mechanism) is the fit, not the other six generalist repositories.
- Persistent identifiers and metadata. Confirm the repository assigns a citable DOI or accession number and captures the metadata fields your DMS Plan commits to — this is one of NIH’s explicit desirable characteristics and is checked at the plan-review and closeout stages.
- Long-term sustainability. NIH also expects a credible plan for the repository’s own long-term viability. GREI’s coordination work is one input into this, but institutions should still confirm current status rather than assume permanence.
For the underlying policy requirements — what a DMS Plan must contain, when it’s due, and how NIH defines scientific data — see CASRAI’s guides to the NIH Data Management and Sharing Policy and the NIH DMS Plan template. For the general distinction this guide relies on, see CASRAI’s dictionary entries on generalist repository and discipline-specific repository, and the guide to research repositories versus archives for the broader vocabulary.
Frequently asked questions
What is GREI?
GREI (Generalist Repository Ecosystem Initiative) is an NIH Office of Data Science Strategy program, launched in 2022, that coordinates shared standards, metadata, metrics, and training across seven established generalist data repositories to better support NIH-funded data sharing.
Which repositories are part of GREI?
Dataverse, Dryad, Figshare, Mendeley Data, Open Science Framework (OSF), Vivli, and Zenodo.
Does NIH require researchers to use a GREI repository?
No. NIH does not designate or require any specific generalist repository, GREI membership included. NIH’s repository-selection guidance places NIH-designated and discipline-specific repositories ahead of any generalist repository, and generalist repositories — GREI members or not — are appropriate only when no suitable discipline-specific option exists.
Is Vivli part of GREI, and how is it different from the other six?
Yes. Vivli is a GREI member, but unlike the other six it is built for controlled, request-and-review access to clinical trial individual participant data rather than open or embargoed deposit. See CASRAI’s guide to how Vivli’s data request and access-review process works for the operational detail.
Who runs GREI?
NIH’s Office of Data Science Strategy (ODSS) funds and coordinates GREI; each of the seven member repositories is independently operated by its own organization.







