The federal “desirable characteristics of trustworthy repositories” checklist that researchers are usually pointed to is a specific, named document: Desirable Characteristics of Data Repositories for Federally Funded Research, published in May 2022 by the Subcommittee on Open Science (SOS) of the National Science and Technology Council (NSTC), under the White House Office of Science and Technology Policy (OSTP). It is not a certification standard, a scoring rubric, or a list a repository can “pass” or “fail” — it is a consensus checklist Federal agencies agreed to use so that data management plan (DMP) guidance about repository selection is consistent across agencies. This guide walks through every characteristic it defines, in the document’s own wording, and explains how research administrators and investigators are actually expected to use it.
What the document is, and where it comes from
The document’s full citation is: The National Science and Technology Council, Desirable Characteristics of Data Repositories for Federally Funded Research, 2022 (DOI: 10.5479/10088/113528). It traces back to the February 2013 OSTP Memorandum on Increasing Access to the Results of Federally Funded Scientific Research, which directed every Federal agency with more than $100 million in annual R&D spending to require funded researchers to submit data management plans. More than 20 agencies subsequently adopted public access and data-sharing policies, but their guidance on which repositories to use was inconsistent. OSTP issued a draft version for public comment in January 2020 (119 responses were received from universities, research consortia, data centers, scientific societies, and government agencies), and the final consensus document was published in May 2022 under Alondra Nelson’s tenure as Acting OSTP Director — the same period that produced the better-known August 2022 “Nelson Memo” on public access to federally funded research.
NIH’s 2023 Data Management and Sharing (DMS) Policy explicitly directs investigators toward repositories that “exemplify” this framework whenever no NIH-designated or discipline-specific repository exists for their data type, which is the main reason the document shows up in day-to-day DMP-writing and repository-selection work.
What it deliberately is not
The document itself is explicit on this point: Federal agencies “have elected not to adopt existing certification criteria, due in part to the cost and complexity of certification processes,” and they state plainly that they “do not plan to use the characteristics in this guidance document to assess, evaluate, or certify the acceptability of a specific data repository, unless otherwise required for a particular agency program, initiative, or funding opportunity.” In other words:
- It is not a certification scheme like CoreTrustSeal or ISO 16363 — no repository is “NSTC-certified,” and there is no application, review, or fee.
- It is not exhaustive. The document says so directly: “the desirable characteristics provided by this guidance document are not intended to be an exhaustive set of features for data repositories; rather they represent general capabilities for researchers, agencies, and institutions to prioritize when selecting repositories.”
- It is designed to be consistent with existing certification criteria — the SOS reviewed criteria from bodies including the International Organization for Standardization and the International Science Council while drafting it, which is why its three organizing themes (Organisational Infrastructure, Digital Object Management, Technology) mirror the structure certification bodies like CoreTrustSeal use.
The 14 desirable characteristics for all repositories
Table 1 of the document lists 14 characteristics organized under three themes. The wording below is drawn directly from the document.
Organizational Infrastructure
| Characteristic | What it requires |
|---|---|
| Free and Easy Access | Broad, equitable, and maximally open access to datasets and their metadata, free of charge, in a timely manner after submission — consistent with legal and policy requirements around privacy, confidentiality, Tribal and national data sovereignty, and protection of sensitive data. |
| Clear Use Guidance | Documentation describing terms of dataset access and use, such as reuse licenses and any requirement for approval by a data use committee. |
| Risk Management | Documented administrative, technical, and physical safeguards to comply with confidentiality, risk-management, and continuous-monitoring requirements for sensitive data. |
| Retention Policy | Published documentation on the repository’s data retention policies. |
| Long-term Organizational Sustainability | A plan for long-term management of data — maintaining integrity, authenticity, and availability of datasets — plus contingency plans covering unforeseen events. |
Digital Object Management
| Characteristic | What it requires |
|---|---|
| Unique Persistent Identifiers | Assigns each dataset a citable, unique persistent identifier (PID/DPI), such as a DOI, that points to a persistent location remaining accessible even if the dataset is later de-accessioned. |
| Metadata | Metadata accompanying every dataset, using schema appropriate to — and ideally widely used across — the communities the repository serves, to enable discovery, reuse, and citation. |
| Curation and Quality Assurance | Provides or facilitates expert curation and quality assurance to improve the accuracy and integrity of datasets and metadata. |
| Broad and Measured Reuse | Metadata describing terms of reuse, plus the ability to measure attribution, citation, and reuse of data (via adequate, openly accessible metadata and unique PIDs). |
| Common Format | Datasets and metadata can be accessed, downloaded, or exported in widely used, preferably non-proprietary, formats consistent with disciplinary standards. |
| Provenance | Mechanisms to record the origin, chain of custody, version control, and any other modifications made to submitted datasets and metadata. |
Technology
| Characteristic | What it requires |
|---|---|
| Authentication | Supports authentication of data submitters, with technical capabilities to associate submitter PIDs with the identifiers assigned to their deposited objects. |
| Long-term Technical Sustainability | A plan for long-term data management built on stable technical infrastructure and funding plans. |
| Security and Integrity | Documented measures meeting established cybersecurity criteria to prevent unauthorized access, modification, or release of data, at a level appropriate to data sensitivity (the document cites the NIST Cybersecurity Framework as an example). |
The 7 additional considerations for repositories storing human data
Table 2 adds characteristics specific to repositories holding de-identified human data, on top of the 14 above, because “re-identification of de-identified human data remains a risk for many datasets.”
| Consideration | What it requires |
|---|---|
| Fidelity to Consent | Documented procedures restricting dataset access and use to what is consistent with participant consent (e.g., use only within a specific disease or condition) and any subsequent changes in consent. |
| Security | Documented approaches — tiered access, credentialing of data users, security safeguards against breaches — to protect human subjects’ data from inappropriate access. |
| Limited Use Compliant | Documented procedures to communicate and enforce data use limitations, such as preventing re-identification or unauthorized re-distribution. |
| Download Control | Controls and audits access to, and download of, datasets. |
| Request Review | An established and transparent process for reviewing data access requests. |
| Plan for Breach | Security measures that include a response plan for detected data breaches. |
| Accountability | Procedures for addressing violations of terms-of-use and data mismanagement. |
How Federal agencies actually intend it to be used
The document names three primary uses, all advisory rather than evaluative:
- Helping Federally funded institutions and investigators identify a repository when the funding agency hasn’t designated one;
- Helping a Federal agency identify or designate repositories for a particular data type; and
- Informing how agencies evaluate the repository-selection sections of submitted data management plans.
It’s this third use that matters most in practice: when an NIH or other Federal reviewer reads the repository-selection portion of a data management plan, they are checking it against something resembling this framework, even though it isn’t scored as a checklist. Because the document also states it may be periodically updated by the SOS “to reflect changing expectations, advances in research and technology, and evolving practices,” treat any specific wording here as accurate to the May 2022 version and re-check the source PDF if citing it in a compliance context years later.
Using it as a practical repository-selection checklist
Because the document doesn’t score or certify, the practical way research administrators use it is as a due-diligence checklist before naming a repository in a DMP:
- Check for a designated repository first. Many agencies specify a repository for particular data types (e.g., genomic data, Arctic research data) — if one exists, the characteristics below don’t govern the choice.
- If no repository is designated, evaluate a candidate data repository against the 14 characteristics above — most of this information is published on a repository’s own policies/about pages.
- Check certification as corroborating evidence, not a substitute. A repository holding CoreTrustSeal certification or similar is a strong (though not equivalent) signal it already satisfies most of these characteristics, since the SOS designed the framework to be consistent with existing certification criteria.
- If the dataset includes human data, also check the 7 additional considerations in Table 2 — consent fidelity, tiered access, and breach response plans in particular.
- Document the match in the DMP’s repository-selection rationale, since this is the section Federal reviewers are most likely to check against this framework.
Frequently asked questions
Is this the same thing as CoreTrustSeal certification?
No. CoreTrustSeal is a formal, fee-based, peer-reviewed certification with 16 numbered requirements that a repository actively applies for and can lose at renewal. The NSTC’s desirable characteristics document is Federal guidance for researchers and agencies to use when selecting a repository — it isn’t a certification a repository holds, and the document explicitly says Federal agencies chose not to build a certification program around it. See our full CoreTrustSeal certification guide for the certification-side comparison.
Does NIH require repositories to meet these characteristics?
NIH’s 2023 Data Management and Sharing Policy directs investigators toward repositories that exemplify the framework specifically when no NIH-designated or discipline-specific repository exists for their data. It is guidance for that selection decision, not a pass/fail compliance gate applied to every repository choice.
Where can I read the original document?
The May 2022 PDF is archived at the Biden White House archive site (bidenwhitehouse.archives.gov) and mirrored by agencies including NASA’s Earthdata program; the recommended citation DOI is 10.5479/10088/113528, registered via the Smithsonian Institution, which provided copyediting and DOI services for the document.
Do all 14 characteristics carry equal weight?
The document doesn’t rank or weight them — it presents all 14 as characteristics to “prioritize when selecting repositories,” organized into three themes rather than a scored rubric. It is explicitly not exhaustive, and agencies may weight or supplement it differently for particular programs.







