Dictionary domainTrack B
Research data infrastructure
Trusted repositories, EOSC, biobanks, data trusts, federated infrastructure.
For implementers
Operational deployment checklist for Research data infrastructure: prerequisites, five deploy steps, integration notes for Pure, Symplectic Elements, Worktribe, DSpace, and more, plus the pitfalls that recur in the field.
Terms in this domain
90 terms
QDR (Qualitative Data Repository)
The Qualitative Data Repository (QDR) is a domain repository dedicated to curating, preserving, and providing access to digital data generated through qualitative and multi-method social science research — interview transcripts, fieldnotes, focus-group recordings, participant-observation records, and other unstructured or semi-structured data types that do not fit the tabular, quantitative deposit model most general-purpose repositories are built around. It is hosted by the Center for Qualitative and Multi-Method Inquiry at Syracuse University's Maxwell School of Citizenship and Public Affairs. A dataset qualifies as a QDR deposit when it is submitted through QDR's own curation and review workflow, receives a persistent identifier (a DataCite DOI), and is described with QDR's qualitative-data-specific metadata and, where the depositing researcher chooses to use it, QDR's Annotation for Transparent Inquiry (ATI) format for linking published claims back to underlying data excerpts.
GBIF (Global Biodiversity Information Facility)
GBIF (the Global Biodiversity Information Facility) is an intergovernmental network and open-access data infrastructure -- coordinated by a Secretariat in Copenhagen and funded by its participating governments and organizations -- that aggregates, indexes, and republishes species-occurrence and specimen records contributed by data-holding institutions worldwide through a global network of national and thematic Participant Nodes. A dataset counts as GBIF-mediated when it has been registered through a GBIF Participant Node (or directly via GBIF.org) and is discoverable, searchable, and downloadable through the GBIF occurrence-search interface and API; the great majority of datasets reach GBIF as Darwin Core Archives (DwC-A), most commonly built and published using GBIF's own Integrated Publishing Toolkit (IPT). GBIF itself does not define a data standard -- it consumes Darwin Core, the pre-existing TDWG biodiversity-data standard -- and is best understood as the largest aggregating platform built on top of that standard, not the standard itself.
GA4GH (Global Alliance for Genomics and Health)
An international not-for-profit alliance, founded in 2013, that develops and maintains technical interoperability specifications and policy frameworks (such as the Data Use Ontology, GA4GH Passport, Beacon, htsget, Crypt4GH, and Phenopackets) enabling genomic and health-related data to be shared responsibly across research and clinical institutions worldwide; GA4GH does not host data itself but defines the standards that repositories, biobanks, and national genomic initiatives implement to interoperate.
Open Database License (ODbL)
The Open Database License (ODbL) is a copyleft license from Open Data Commons (a project of the Open Knowledge Foundation), version 1.0 published June 2009, that grants reuse rights to a database subject to attribution and share-alike conditions, drafted specifically around database rights and database structure/contents rather than adapted from a creative-works copyright license.
FAIR Data Principles
The FAIR data principles are a set of four foundational goals — Findable, Accessible, Interoperable, and Reusable — for how research data and its metadata should be described, deposited, and structured so that both humans and machines can locate, retrieve, and reuse it with minimal manual effort. FAIR is a set of guiding goals for data stewardship, not a certification, a checklist with one correct implementation, or a data-sharing mandate in itself: a dataset is 'FAIR' to the degree its metadata and deposit environment satisfy the fifteen sub-principles across the four categories, and different repositories, disciplines, and data types satisfy them by different concrete means.
Metadata
Structured information that describes, identifies, or explains a research output — most often a dataset — independent of the content itself, so that the output can be found, correctly interpreted, cited, and reused without the data (or document, sample, or instrument) having to be opened or examined directly. In research data management, metadata is what a repository, catalog, or search index actually reads to determine whether a record matches a query; the underlying data file is opaque to that process until metadata points to it.
ONIX (ONline Information eXchange)
An XML-based family of metadata standards, maintained internationally by EDItEUR (with regional groups Book Industry Communication/BIC in the UK and BISG in the US), used to exchange structured product and rights information across the publishing supply chain. A metadata record qualifies as ONIX when it validates against one of EDItEUR's published ONIX XML schemas/DTDs (the current major version for book-trade metadata is ONIX for Books 3.0, extended by 3.1) and carries the standard's defined composite structure, code lists, and identifiers rather than a proprietary or free-text feed.
Microsoft Academic Graph (MAG)
A large-scale, free scholarly knowledge graph built and maintained by Microsoft Research — publications, authors, affiliations, venues, and a machine-generated fields-of-study taxonomy, linked by citation and authorship edges, and distributed via the Microsoft Academic website, the Microsoft Academic Knowledge API, and periodic bulk data dumps — that Microsoft fully retired on December 31, 2021, and no longer maintains, updates, or serves in any form.
OAIS Reference Model (ISO 14721)
A conceptual framework, standardised as ISO 14721, that defines the functions, information packages, and terminology an archive needs to preserve digital (or physical) information and keep it accessible to a defined community over the long term, independent of the specific technology used to implement it.
Research Data
Research data is the recorded, factual material generated or collected in the course of a research project that is needed to validate, reproduce, or build on that project's findings — raw measurements, observations, instrument readings, survey responses, code outputs, and other structured or unstructured records, regardless of medium or discipline. In U.S. federally funded research specifically, the term carries a narrower, load-bearing regulatory meaning under 2 CFR § 200.315(e)(3) (the Uniform Guidance, which consolidated and replaced the older OMB Circular A-110 in 2014): research data means <em>'the recorded factual material commonly accepted in the scientific community as necessary to validate research findings.'</em> That specific definition is what determines what a federal award recipient must make available in response to a Freedom of Information Act (FOIA) request, and it is the baseline many institutions use to scope what a <a href='/dictionary/term/data-management-plan-dmp'>data management plan (DMP)</a> is actually obligated to describe. Distinguishing 'research data' in this operational sense from the much larger set of everything a researcher produces during a project — drafts, correspondence, physical samples, unfinished analysis — is the first practical step in scoping a DMP, a repository deposit, or a records-retention schedule, and underpins <a href='/pillar/rdm'>research data management (RDM)</a> practice generally.
Administrative Metadata
Administrative metadata is the category of metadata used to manage a resource — as distinct from metadata used to describe or discover it. It typically covers four practical areas: rights and permissions (licensing, access restrictions), provenance and acquisition history (who deposited it, when, and how), technical/file-format details (file type, software/equipment used, checksums), and access-control information. NISO's standard three-way split names it alongside descriptive and structural metadata, and treats technical and preservation metadata as administrative subtypes — though usage varies slightly by standard: METS, for example, models technical, rights, provenance, and preservation information together under a single amdSec (administrative metadata section), sitting alongside — not nested inside — its descriptive metadata section (dmdSec).
NIH Genomic Data Sharing (GDS) Policy
A study falls under NIH's Genomic Data Sharing (GDS) Policy (2014) when it is NIH-funded and generates large-scale human or non-human genomic data (GWAS, SNP arrays, WGS/WES, transcriptomic, epigenomic, metagenomic, or gene-expression data). Covered human-data studies must use GDS-specific prospective informed consent language, obtain an Institutional Certification from the awardee institution (via its Signing Official, in coordination with the IRB) confirming consistency with participant consent, and submit data to an NIH-designated controlled-access repository -- dbGaP for most human genomic data -- under a Data Use Certification Agreement. The GDS Policy predates and operates alongside, not inside, the broader 2023 NIH Data Management and Sharing (DMS) Policy.
Technical Metadata
Technical metadata is the category of metadata that documents the digital characteristics of a resource needed to render, verify, and process it correctly over time: file format and version, the hardware and software environment used to create or read it, technical characteristics (size, encoding, resolution, compression), and a record of any technical processing or transformation events (format migration, compression, checksum regeneration) applied since creation. It answers 'what does a machine need to know to open, verify, and correctly interpret this file,' as distinct from what a resource is about (descriptive metadata) or who may use it and under what terms (administrative metadata, in the narrower rights/provenance/access-control sense). Some frameworks, including NISO's influential three-way split, treat technical metadata as a subtype nested under the broader administrative metadata category; others, including the METS schema's own internal structure, give it a distinct named subsection (techMD) alongside rights and provenance subsections rather than folding it into a single undifferentiated administrative bucket. Both framings agree on the underlying content -- format, environment, and fixity information -- they differ only in where they draw the taxonomy line.
Metadata Format
A metadata format (also called a serialization format) is the concrete syntactic encoding used to write metadata values out as a machine-readable byte stream for storage, exchange, or transmission — for example XML, JSON, RDF/Turtle, or YAML. It answers 'how is this metadata physically written down,' which is a separate question from what a metadata schema answers: 'which fields exist and what do they mean.' A schema (Dublin Core, DataCite, Darwin Core, PROV-O) defines the field set and semantics; a format defines the byte-level syntax those fields get poured into. The same schema is routinely available in more than one format, and the same format is routinely used to carry many unrelated schemas — so 'metadata format' and 'metadata schema' are not interchangeable, even though the two terms get conflated constantly in casual usage.
Globus
A non-profit, hosted research data management service - built on GridFTP-based transfer technology and operated by the University of Chicago with Argonne National Laboratory - that provides high-performance, fault-tolerant file transfer, institutional 'endpoints' (via Globus Connect Server/Personal), and controlled sharing between research storage systems, HPC centers, and repositories.
Australian Research Data Commons (ARDC)
The national research data infrastructure provider for Australia: a not-for-profit company, funded through the Australian Government's National Collaborative Research Infrastructure Strategy (NCRIS), that builds, operates, and stewards shared digital research infrastructure — including platforms, skills programs, and persistent-identifier services such as RAiD — for use across the Australian university and public-research sector rather than by any single institution.
ISO 16363 (Formal-Level Certification)
The Formal level of the three-tier trustworthy-repository certification framework: a repository is ISO 16363-certified only when an accredited external auditor (typically PTAB, accredited under ISO 17021/16919) has conducted a full audit against ISO 16363's evidence-based metrics and formally certified the result — distinct from CoreTrustSeal's and the nestor Seal's self-assessment models. Full title: Space data and information transfer systems — Audit and certification of trustworthy digital repositories. Descends from the 2007 TRAC (Trustworthy Repositories Audit & Certification) checklist published by RLG/NARA, formalised by CCSDS as ISO 16363:2012 and subsequently revised.
Secure Data Enclave
A controlled-access computing environment in which approved researchers analyze restricted or sensitive data in place, without the ability to download, copy, or otherwise export the raw underlying records. Access is typically granted only after training, credentialing, and a signed data use agreement, computation happens on infrastructure the data steward controls (an on-site terminal room, a remote-access session, or an isolated cloud workspace), and only aggregate or disclosure-reviewed outputs — tables, statistics, model results — are allowed to leave the environment. The defining feature is not where the data sits but that raw data never crosses the boundary; only vetted outputs do.
Repository Citation
A repository citation is the reference used to attribute and locate a dataset (or other research output such as software or a physical sample) deposited in a data repository, built around a persistent identifier — typically a DOI minted via DataCite — rather than around the journal, volume, and page numbers used for an article citation. A citation counts as a repository citation when it points directly to the deposited data object itself (resolving through its own persistent identifier to the repository landing page or the data), not to a journal article, data paper, or other publication that merely describes or analyzes that data.
Concordat on Open Research Data
The Concordat on Open Research Data is a UK sector-wide policy framework, published 28 July 2016 by the Higher Education Funding Council for England (HEFCE), Research Councils UK (RCUK, since absorbed into UK Research and Innovation, UKRI), Universities UK, and the Wellcome Trust, setting out ten principles for how researchers, research organisations, and funders should approach making research data openly available. It is not a regulation, mandate, or legal instrument: an institution, funder, or research group operationalises the Concordat when its research data policy, Data Management Plan (DMP) guidance, or grant terms explicitly build in the ten principles — for example, requiring a documented, case-by-case justification for any restriction on data openness (Principle 2), recognising a researcher's right to a reasonable first-use period before data must be shared (Principle 4), or expecting data underlying a publication to be accessible, citable via a persistent identifier, and retained for at least ten years from the publication date (Principle 8).
Bring Your Own Device (BYOD) for Research Data Collection
In research data collection, BYOD (bring your own device) is a practice in which participants or researchers use their own personally-owned smartphone, tablet, or wearable -- rather than a device supplied, configured, and owned by the study -- to run a data-collection app or sensor. What makes an instance BYOD is device ownership and control: the hardware belongs to the participant or researcher, not the study team, so the study cannot fully standardize, lock down, or reclaim it, and must build its data-integrity and security controls around that constraint.
PREMIS
A resource's metadata qualifies as PREMIS metadata when it is structured according to the PREMIS Data Dictionary for Preservation Metadata (currently version 3.0, maintained by the Library of Congress), populating semantic units under one or more of PREMIS's four in-scope entities — Object, Event, Agent, and Rights — to document what a digital object is, what has happened to it over time, who or what acted on it, and what permissions govern its preservation. PREMIS is a data dictionary, not a file format or package standard in the way METS is; it is most commonly implemented as an XML schema (premis.xsd) embedded inside a METS document's administrative metadata section, or serialized natively inside a repository's own preservation system.
DDI Codebook
A metadata document is an instance of DDI Codebook (DDI-C) when it is a well-formed XML file rooted at the <codeBook> element and containing the DDI Alliance's standard sections for describing a single dataset: a study description (title, authors, abstract, methodology, sampling procedure), a file description of the physical data file, and a data description that documents each variable individually, including its name, label, question text, value labels, and universe. It is maintained by the DDI Alliance and versioned (current releases in the 2.x line, e.g. 2.5/2.6); it is the single-dataset-focused branch of DDI, distinct from the multi-wave/lifecycle-tracking DDI Lifecycle (DDI 3.x) specification.
W3C DCAT (Data Catalog Vocabulary)
DCAT (Data Catalog Vocabulary) is a W3C Recommendation: an RDF vocabulary, published under the namespace http://www.w3.org/ns/dcat#, for describing datasets, dataset series, distributions, and data services within a data catalog so that catalogs published by different organizations can be harvested, aggregated, and searched interoperably. A metadata record qualifies as DCAT (rather than a merely similar, ad hoc catalog schema) when its fields map to DCAT's own RDF classes and properties -- dcat:Catalog, dcat:Dataset, dcat:Distribution, dcat:DataService, dcat:DatasetSeries, and dcat:CatalogRecord, built on top of Dublin Core, FOAF, SKOS, and PROV-O terms rather than a locally invented field set -- and is expressed (or losslessly convertible to) RDF, typically serialized as JSON-LD, Turtle, or RDF/XML.
ISO 23081 (Records Management Processes — Metadata for Records)
ISO 23081 (Information and documentation — Records management processes — Metadata for records) is the ISO standard that defines the principles and framework governing the metadata records systems must capture so that records can be managed, retrieved, and trusted as evidence throughout their lifecycle. A metadata scheme qualifies as ISO 23081-aligned when it goes beyond simple descriptive metadata (title, creator, date) to capture recordkeeping context: the business process or mandate that created the record, the agents responsible for it, its relationships to other records and aggregations, and its retention/disposition authority. The standard is published in three parts: Part 1 sets out the principles underpinning records metadata; Part 2 gives conceptual and implementation guidance for building a metadata schema consistent with Part 1; Part 3, published as a Technical Report, provides a self-assessment method for evaluating an existing metadata schema against the standard. ISO 23081 was developed by ISO/TC 46/SC 11 as a companion to ISO 15489 (the core international records management standard), translating 15489's principles into a concrete metadata model that records-management systems, archives, and institutional repositories can actually implement.
ANSI/NISO Z39.19 (Controlled Vocabularies and Thesauri)
A controlled vocabulary or thesaurus is constructed to ANSI/NISO Z39.19-2005 (R2010), "Guidelines for the Construction, Format, and Management of Monolingual Controlled Vocabularies," when it follows the standard's rules for term selection and form (preferring noun phrases, singular/plural conventions, avoiding ambiguous homographs without a qualifier), and when it displays the three relationship types Z39.19 defines between terms: equivalence relationships (USE/UF, linking a non-preferred synonym to its preferred term), hierarchical relationships (BT/NT, Broader Term/Narrower Term, for genus-species, whole-part, or instance relationships), and associative relationships (RT, Related Term, for terms that are conceptually related but not synonymous or hierarchical). The standard covers several vocabulary types along a spectrum of structural complexity -- authority lists, synonym rings, taxonomies, and full thesauri -- and specifies how each should be tested, formatted, and maintained over time (adding, deprecating, and merging terms as usage and the underlying domain evolve).
Data Curation Network (DCN)
The Data Curation Network (DCN) is a membership consortium of academic and non-profit research-data repositories that pools trained data curators across member institutions so that any participating library can call on specialist review for a dataset, rather than each institution having to hire full curation expertise for every data type it receives. It is distinct from 'data curation' the generic activity and 'data curator' the generic role: DCN is a specific, named organization, hosted administratively by the University of Minnesota Libraries, that operationalizes those generic concepts as a shared-staffing service across its member institutions.
Open Science Framework (OSF)
A free, open-source web platform operated by the nonprofit Center for Open Science (COS) that lets a research team manage a project's full lifecycle from one persistent, citable project page: file storage and version control, collaborator permissions, preregistration of hypotheses and analysis plans, and public archiving with a DOI. 'Open Science Framework' is often shortened or misremembered as 'Open Science Foundation' — the organization that builds and runs OSF is the Center for Open Science, not an 'Open Science Foundation'; no organization by that name operates this infrastructure.
Metadata Schema
A defined, published set of descriptive fields (elements or properties) — each with a name, definition, expected data type or controlled vocabulary, and obligation/cardinality rules — used to describe a resource consistently enough for people and systems to find, cite, exchange, and reuse it. Distinct from a metadata record, which is one specific instance of data filled into that specification.
Electronic Medical Record (EMR)
An electronic medical record (EMR) is the digital equivalent of a single healthcare provider's paper chart: a patient's medical and treatment history as captured, stored, and used within one clinical practice or healthcare organization. An artifact is an EMR (rather than an EHR) when its data model, access controls, and interoperability are scoped to a single organization's internal clinical workflow, not designed for structured exchange with external providers, payers, or research systems -- even if the underlying software vendor also sells EHR-branded products elsewhere.
Electronic Data Capture (EDC)
Electronic Data Capture (EDC) is the general category of software systems that collect clinical or research study data directly in electronic form -- via electronic case report forms (eCRFs) -- at the point of collection, in place of paper-based data collection. A system counts as EDC when it provides study-specific eCRFs, direct point-of-collection data entry, field-level validation and query management, and an audit trail sufficient to make the electronic record the study's data of record. EDC is the site-facing data-entry layer within the broader discipline of clinical data management; a full Clinical Data Management System (CDMS) is the broader term covering EDC plus surrounding query-management, coding, and database-lock workflow. EDC is distinct from a Clinical Trial Management System (CTMS), which manages trial operations rather than the research data itself. Both non-profit/academic platforms (e.g., REDCap) and commercial, enterprise-scale platforms (e.g., Medidata Rave EDC, Oracle Clinical One) fall within the EDC category.
Electronic Theses and Dissertations (ETD)
A thesis or dissertation prepared, submitted, and archived in digital form (typically a PDF plus any supplementary files) rather than only as a bound print copy, submitted to satisfy a degree requirement. A work qualifies as an ETD once digital deposit — most often to an institutional repository, and often onward to a national or international aggregator such as NDLTD — is the accepted or required submission format, even if the institution also retains a print copy alongside it.
Data Ownership
The allocation of legal title to and decision-making control over a research dataset. Because raw data is rarely copyrightable and no single U.S. statute assigns default ownership of research data, the operative answer in practice comes from a combination of institutional policy, the funder's terms and conditions, the Data Management Plan, and any executed data sharing/use agreement -- not from a single ownership doctrine. Distinguish from data stewardship/custodianship, which describes accountability for a dataset's day-to-day management regardless of who holds title.
protocols.io (concept)
protocols.io is a free, open platform (owned by Springer Nature since 2023) for developing, versioning, and sharing detailed step-by-step research method protocols across life-science, physical-science, and computational disciplines. Each published protocol receives a DataCite DOI, so a specific, citable version of a method persists even as the protocol is later revised, forked, or optimized by other labs.
REDCap (Research Electronic Data Capture)
REDCap (Research Electronic Data Capture) is a secure, metadata-driven, web-based software platform for building surveys and case-report-form databases for research, originated at Vanderbilt University in 2004. An instance only counts as REDCap if it runs the actual Vanderbilt-distributed software, licensed at no cost to non-profit institutions through the REDCap Consortium and installed/administered by that partner institution -- there is no direct commercial purchase path or self-service individual signup. Typical uses include participant surveys, longitudinal case report forms, and multi-site clinical data capture, frequently chosen specifically because the platform supports HIPAA, FDA 21 CFR Part 11, FISMA, and GDPR compliance controls at the institutional level.
3-2-1 Backup Strategy
A data-protection rule requiring that any dataset worth protecting exist as at least three total copies, stored across at least two different storage media or systems, with at least one of those copies kept at a location or system genuinely separate from where the working copy lives. All three conditions must hold simultaneously -- extra copies on the same medium, or an offsite copy with no independent local redundancy, do not satisfy the rule on their own.
Dublin Core
Dublin Core is the cross-domain metadata vocabulary, governed by the Dublin Core Metadata Initiative (DCMI), built around fifteen core resource-description elements (Title, Creator, Subject, Description, Publisher, Contributor, Date, Type, Format, Identifier, Source, Language, Relation, Coverage, Rights) formally standardized as ISO 15836, ANSI/NISO Z39.85, and IETF RFC 5013 (Simple Dublin Core), plus the larger DCMI Metadata Terms ("dcterms:") vocabulary that adds element refinements, encoding schemes, and additional properties/classes for more precise description (Qualified Dublin Core). A metadata record is a Dublin Core record when its fields map to this DCMI-maintained element set or the extended dcterms vocabulary, rather than to an unrelated, locally invented field set or an unrelated domain-specific schema.
Data Curator
A data curator is an institutional role — usually in a library, research data service, or research-computing unit — that prepares research data for deposit, applies metadata standards, checks FAIR readiness, and advises researchers on repository selection and data management plans, distinct from owning the underlying research itself.
Data Curation
The active, ongoing management of research data across its lifecycle -- organizing files into a coherent structure, describing them with standardized metadata, validating and cleaning values, and taking preservation actions such as format migration and fixity checking -- carried out specifically so the data remains findable, interpretable, and reusable by someone other than its original creator, long after the project that produced it ends. Curation is a defined set of actions applied to data over time; it is distinct from simply storing a copy of it.
dbGaP (Database of Genotypes and Phenotypes)
The NIH/NCBI repository for genotype-phenotype study data (GWAS results, sequencing/omics data, linked phenotype data), split into an open-access tier (study documentation, summary statistics) and a controlled-access tier (de-identified individual-level genotype and phenotype records) gated by Data Access Committee review of a Data Access Request -- distinct from a general sequence archive because the individual-level linkage, not just the data type, is what triggers controlled access.
PDF/A (ISO 19005)
A file is PDF/A-conformant only when it validates against one of the four parts of ISO 19005: every font it uses is embedded rather than referenced, color is specified in a device-independent way, and the file contains no encryption, no JavaScript or executable content, no audio or video, and no reference to external content the renderer would need to fetch. These restrictions make the file self-contained and reproducible without depending on the specific software environment that created it.
PROV-O (Provenance Ontology)
PROV-O is the W3C's OWL2 ontology (a W3C Recommendation since 30 April 2013) for expressing data provenance as machine-readable RDF: it defines Entity (a thing with fixed aspects), Activity (something that acts upon or generates entities over time), and Agent (something responsible for an activity or entity), connected by properties such as wasGeneratedBy, used, wasAssociatedWith, and wasDerivedFrom.
Data Repository
A system or platform dedicated to the long-term storage, curation, preservation, and dissemination of research data, providing persistent identifiers, standardized metadata, and defined access/reuse terms -- distinct from general-purpose file storage (a network drive or cloud folder), which does not guarantee a dataset stays findable, citable, or usable once active work on it ends.
Core metadata
The minimal, standard set of descriptive fields — typically title, creator, date, persistent identifier, subject/keywords, format, and rights — that a dataset or research output must carry in order to be findable and citable, independent of any richer, discipline-specific metadata layered on top of it.
Darwin Core
A data standard ratified by TDWG (Biodiversity Information Standards) — a glossary of defined terms organized into a small set of classes (Occurrence, Taxon, Event, Location, MaterialEntity, Identification, and others) — used to structure and exchange biodiversity data documenting the occurrence of organisms and the specimens or observations that record them. A dataset counts as Darwin Core data when its fields are actually mapped to defined Darwin Core terms (e.g. dwc:scientificName, dwc:eventDate, dwc:decimalLatitude), not merely because it describes species or specimens in some other format.
Extended level certification (trustworthy repository)
The middle tier of a three-level trust framework for research-data repositories, positioned above core level (CoreTrustSeal, a peer-reviewed self-assessment against 16 requirements) and below formal level (ISO 16363, a full external audit). A repository holds extended level certification when it has been awarded the nestor Seal for Trustworthy Digital Archives — a plausibility-checked self-assessment against the 34 criteria of the German standard DIN 31644, administered by the nestor competence network at the Deutsche Nationalbibliothek. It is more rigorous than a core-level self-assessment (more criteria, external plausibility review of the evidence submitted) but, unlike formal level, does not involve an accredited external auditor conducting an on-site or fully independent audit.
Aggregator service
A service that harvests, harmonises, and re-exposes metadata and (sometimes) content from many upstream sources, providing a unified search, browse, or query interface across the aggregated corpus; canonical examples include OpenAIRE, BASE, CORE, and OpenAlex.
Data safe haven
A secure data-handling environment that allows controlled, audited access to sensitive datasets for approved research, applying technical, physical, and procedural safeguards; effectively a synonym for trusted research environment (TRE) in much current usage, though the term has older roots in NHS information governance.
Five Safes framework
A framework for the safe use of sensitive data in research, articulated by the UK Office for National Statistics, that organises controls under five dimensions: Safe People, Safe Projects, Safe Settings, Safe Data, and Safe Outputs.
Trusted research environment
A secure computing environment — typically delivered as a remote-access workspace with controlled inbound/outbound data flows — that allows accredited researchers to analyse sensitive data in situ without exporting the data, supporting privacy-preserving secondary research use.
Sensitive-data repository
A repository specifically designed to hold sensitive research data — typically personal data, health data, criminal-justice data, commercially-confidential data, or culturally-sensitive Indigenous data — with enhanced access controls, audit logging, contractual access conditions, and (often) a secure analysis environment.
Dataset landing page
The human-readable web page that a dataset's persistent identifier (typically a DataCite DOI) resolves to, presenting the dataset's title, creators, description, identifiers, dates, version history, related works, access conditions, and a link to download or request the data.
Joint Declaration of Data Citation Principles
The 2014 statement produced by Force11's Data Citation Synthesis Group, signed by a wide community of publishers, funders, repositories, and infrastructure providers, that articulates eight principles for the citation of research data in scholarly communication.
Data citation principle
Any of the eight principles articulated in the Joint Declaration of Data Citation Principles (Force11, 2014) covering importance, credit and attribution, evidence, unique identification, access, persistence, specificity and verifiability, and interoperability and flexibility of data citations in scholarly communication.
Data publication platform
A platform that supports the publication of research data as a citable artefact — assigning a persistent identifier, presenting a landing page, and applying review, curation, or peer-review processes — distinct from purely depositional storage.
Domain repository
Synonym for discipline-specific repository: a repository whose scope is a particular research domain (or domain-sub-area), with curation practices and metadata tailored to that domain.
Generalist repository
A repository that accepts research outputs from any discipline, applying domain-agnostic curation and discovery, and serving as a deposit destination for outputs that have no natural discipline-specific home or whose authors prefer a single multidisciplinary venue.
Discipline-specific repository
A repository whose scope is bounded to a particular research discipline or sub-discipline, with curation practices, metadata schemas, and community standards tailored to that domain's data types, terminologies, and norms.
FAIRsharing (concept)
A curated, community-driven registry of databases, standards (metadata, identifiers, formats, terminologies), and data policies relevant to research data, maintained at the University of Oxford with linkage to funders, journals, and standards organisations.
Re3data (concept)
Registry of Research Data Repositories: a global registry, operated by DataCite and partner institutions, that lists research data repositories worldwide with descriptive metadata about their disciplines, content types, access conditions, and policies, helping researchers locate suitable repositories for deposit and discovery.
UK Data Service (concept)
A UK ESRC-funded data infrastructure that holds, curates, and provides access to social, economic, and population data resources for research, learning, and policy, comprising the UK Data Archive at the University of Essex and partner institutions.
ICPSR (concept)
Inter-university Consortium for Political and Social Research: a consortium-membership-funded data archive based at the University of Michigan that holds and curates over 10,000 social-science research datasets, providing access to member institutions worldwide.
Harvard Dataverse (concept)
A free research-data repository operated by Harvard University on the open-source Dataverse software platform, accepting datasets from researchers worldwide, minting DataCite DOIs, and serving as the flagship instance of the global Dataverse network.
Dryad (concept)
A non-profit generalist research data repository operated by Dryad Data Inc. (in partnership with the California Digital Library) that publishes peer-reviewed-paper-linked datasets, mints DataCite DOIs, and applies curation review before publication.
Figshare (concept)
A commercial generalist research repository operated by Digital Science that accepts datasets, figures, presentations, papers, software, and other research artefacts, minting DataCite DOIs and offering institutional-branded instances ('Figshare for Institutions') alongside the public service.
Zenodo (concept)
A free generalist research repository operated by CERN and developed under OpenAIRE that accepts deposits of datasets, software, publications, presentations, posters, and other research artefacts, minting DataCite DOIs and providing free preservation up to a per-record size limit.
GitHub mirror
A copy of a Git repository (or set of repositories) hosted on GitHub that tracks an upstream source repository elsewhere, typically maintained for redundancy, visibility, or community-engagement reasons rather than as the canonical primary copy.
Software Heritage archive
A non-profit international initiative based at Inria that systematically crawls, archives, and preserves the world's publicly available source code, including its full version-control history, and issues persistent identifiers (Software Hash Identifiers, SWHIDs) to every archived artefact.
Code repository
A version-controlled storage location for source code, typically operated on top of a distributed version-control system such as Git, exposing the code's full revision history, branches, tags, and (often) collaboration features such as issues, pull requests, and code review.
Tissue bank
A specific kind of biobank focused on the collection, processing, storage, and distribution of human tissue samples (typically solid tissue specimens from surgical or post-mortem sources), governed under tissue-banking regulation in the relevant jurisdiction.
Sample repository
A repository for physical research samples — geological, environmental, biological, or material — that catalogues, stores, and provides access to samples for downstream analysis, often issuing persistent identifiers (IGSN, DataCite DOI) for citation and provenance tracking.
Biorepository
A facility or organisation that collects, processes, stores, and distributes biological materials and their associated data for research, encompassing both human and non-human samples, distinguished from a 'biobank' by usage in some communities to denote broader scope or specific research projects.
Biobank
An organised collection of biological samples (typically human samples such as blood, tissue, DNA, urine) together with their associated clinical, demographic, and lifestyle data, governed for use in biomedical research.
National data infrastructure
A coordinated, nationally-scoped programme and set of services for the storage, sharing, and reuse of research data within a country, typically combining funding policy, technical infrastructure (repositories, compute, federation), training, and governance.
Data hub
A central node in a data ecosystem that aggregates, harmonises, and brokers access to data from multiple upstream sources, exposing the harmonised data to downstream consumers via curated APIs, query interfaces, or download endpoints.
Federated data infrastructure
A data infrastructure in which data, services, and access controls remain distributed across multiple independent nodes (typically operated by different organisations) but are made discoverable, queryable, and usable as a unified resource through shared protocols, vocabularies, and identity-federation.
Data warehouse
A central repository of structured data, integrated from multiple operational sources, modelled for analytical querying (typically with a star or snowflake schema), and optimised for read-heavy workloads supporting reporting and decision-making.
Data lake
A storage repository that holds large volumes of structured, semi-structured, and unstructured data in their native formats, deferring schema-on-write requirements so that data can be ingested cheaply and only structured at the time of read or analysis.
Data commons
A shared data resource — often combined with shared computing and analysis tools — governed by a community under defined access and contribution rules, designed to enable many users to use and add to the resource for collective benefit.
Data trust
A legal and organisational structure in which a fiduciary intermediary holds, governs, and brokers access to a body of data on behalf of its contributors and beneficiaries, applying agreed terms of access, use, and accountability.
World Data System certification
Historic certification programme of ICSU's World Data System (WDS) under which scientific data centres in geosciences and related fields were certified as trustworthy; merged with the Data Seal of Approval in 2017 to form CoreTrustSeal.
CoreTrustSeal
A community-based, non-profit certification scheme for trustworthy data repositories, operated by the CoreTrustSeal Foundation, awarded against 16 published requirements covering organisational infrastructure, digital object management, and technical infrastructure.
Trusted digital repository
A digital repository whose mission, governance, technical infrastructure, and procedures have been independently assessed against a recognised standard (e.g. CoreTrustSeal, nestor seal, ISO 16363) and judged trustworthy to preserve digital content over the long term.
Electronic Lab Notebook (ELN)
An electronic lab notebook (ELN) is software that replaces the paper laboratory notebook as the primary record of experimental work. An entry qualifies as ELN record-keeping when it is timestamped and attributed at the moment of creation, when edits are versioned rather than overwritten (so a full history remains retrievable), and when the record is structured enough to be searched, exported in a non-proprietary format, and -- increasingly -- assigned a DOI and deposited as a citable research output.
Subject repository
A repository the contents of which are connected purely by their discipline, rather than by other factors such as their institutional affiliation (see Institutional Repository)
Researcher webpage
A webpage featuring a researcher's profile, which possibly may also provide links to their publications.
Repository
Repositories preserve, manage, and provide access to many types of digital materials in a variety of formats.
Open archive
A repository that is compliant with the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) and therefore facilitates the sharing of metadata for a variety of purposes, most notably the compilation tasks performed by aggregator databases.
Institutional webpage
A webpage that is associated with the institution at which the author is employed.
Institutional repository
An online, digital collection of research outputs (see Repository) that are connected by their affiliation with a specific institution. Institutional repositories are most commonly associated with universities and other academic organisations, and so the contents of a single institutional repository may therefore cover a range of disciplines. An institutional repository may often be managed as part of a wider suite of services supporting scholarly communication, Open Access and Open Education.







