Skip to main content
v2026.11,610 entries · CC-BY 4.0

RDM glossary · CC-BY 4.0

Research Data Management Glossary

A research data management glossary is a curated set of defined terms for how research data is planned, stored, shared, preserved, and cited. This RDM glossary is the CASRAI-curated set: 336 research data management terms with stable URIs, grouped by sub-theme and linked to their full definitions in the CASRAI Dictionary.

What is in this glossary

RDM terms
336
Sub-themes
6
Source domains
5
Licence
CC-BY
  • · Part of the CASRAI Dictionary (the full term set, CC-BY 4.0)
  • · Stable URI per entry
  • · Schema.org DefinedTermSet markup
  • · Aligned with FAIR, RDA, and DataCite vocabularies

About this glossary

What research data management terms cover

Research data management (RDM) is the set of practices, stewarded within the wider CASRAI standards portfolio, for handling the data produced by research throughout its lifecycle — from the planning captured in a data management plan, through collection and documentation, to storage, sharing, preservation, and citation. A shared vocabulary matters here more than almost anywhere else in research administration: funders mandate DMPs, repositories enforce metadata schemas, and persistent identifiers stitch outputs back to the people and projects that produced them. When everyone references the same definition for a trusted research environment, a machine-actionable DMP, or a DOI, systems exchange records without manual mapping and reporting becomes auditable across institutions.

This RDM glossary curates 336 terms from five thematic domains of the CASRAI Dictionary — research data infrastructure, machine-actionable data management plans, the persistent identifier ecosystem, research-information systems, and reproducibility and computational research. Each entry below links to its full definition, picklists, related terms, and structured data. The glossary aligns with the FAIR principles, the RDA DMP Common Standard, and the DataCite and CrossRef metadata schemas, and is published under CC-BY 4.0 for unrestricted reuse — including by controlled-vocabulary services such as the UN FAO AGROVOC thesaurus that reference CASRAI definitions.

The discipline this glossary models — one operational definition per term, a stable URI, primary-source verification, and Schema.org DefinedTermSet markup — travels well beyond research administration. Organisations that publish reviewed reference content increasingly pair the same structure with named, role-attributed authorship to make their expertise machine-verifiable. A worked example from healthcare commerce: LAC Health medical-equipment standards glossary documents FDA, ISO, IEC, and ASTM procurement requirements as a DefinedTermSet written from each issuing authority’s primary text, and declares its editorial contributor roles using the CRediT taxonomy’s canonical role URIs, tying the named author to a single cross-site entity via sameAs. For a standards body, that is the adoption pattern working exactly as intended: the controlled-vocabulary and attribution models leaving academia and structuring commercial reference content the same way they structure a CRIS record or a data-management plan.

Part of the CASRAI Dictionary (the full term set, CC-BY 4.0). For institution-level adoption guidance see /for-institutions; for reproducibility standards see /standards/reproducibility.

Data management and governance

Core concepts for stewarding research data across its lifecycle, from secure handling environments to governance constructs.

For the applied guidance behind these terms, see data privacy, de-identification and data-sharing agreements.

16 terms

Aggregator service
A service that harvests, harmonises, and re-exposes metadata and (sometimes) content from many upstream sources, providing a unified search, browse, or query interface across the aggregated corpus; canonical examples include OpenAIRE, BASE, CORE, and OpenAlex.
Data commons
A shared data resource — often combined with shared computing and analysis tools — governed by a community under defined access and contribution rules, designed to enable many users to use and add to the resource for collective benefit.
Data hub
A central node in a data ecosystem that aggregates, harmonises, and brokers access to data from multiple upstream sources, exposing the harmonised data to downstream consumers via curated APIs, query interfaces, or download endpoints.
Data lake
A storage repository that holds large volumes of structured, semi-structured, and unstructured data in their native formats, deferring schema-on-write requirements so that data can be ingested cheaply and only structured at the time of read or analysis.
Data safe haven
A secure data-handling environment that allows controlled, audited access to sensitive datasets for approved research, applying technical, physical, and procedural safeguards; effectively a synonym for trusted research environment (TRE) in much current usage, though the term has older roots in NHS information governance.
Data steward role (in DMP)
A named individual or role accountable for the day-to-day execution of a DMP's commitments, typically distinct from the principal investigator and reported in the DMP's contributor metadata.
Data trust
A legal and organisational structure in which a fiduciary intermediary holds, governs, and brokers access to a body of data on behalf of its contributors and beneficiaries, applying agreed terms of access, use, and accountability.
Data warehouse
A central repository of structured data, integrated from multiple operational sources, modelled for analytical querying (typically with a star or snowflake schema), and optimised for read-heavy workloads supporting reporting and decision-making.
Federated data infrastructure
A data infrastructure in which data, services, and access controls remain distributed across multiple independent nodes (typically operated by different organisations) but are made discoverable, queryable, and usable as a unified resource through shared protocols, vocabularies, and identity-federation.
Five Safes framework
A framework for the safe use of sensitive data in research, articulated by the UK Office for National Statistics, that organises controls under five dimensions: Safe People, Safe Projects, Safe Settings, Safe Data, and Safe Outputs.
National data infrastructure
A coordinated, nationally-scoped programme and set of services for the storage, sharing, and reuse of research data within a country, typically combining funding policy, technical infrastructure (repositories, compute, federation), training, and governance.
Preservation commitment (in DMP)
A statement in a DMP specifying which datasets will be preserved beyond project end, in which repository, for how long, and under what conditions.
Retention period (in DMP)
The minimum duration for which a specified dataset will be retained after project end, expressed as a calendar period and recorded as part of the DMP's preservation commitments.
Sensitive-data handling (in DMP)
The section of a DMP that documents the categories of sensitive data involved (personal, special category, commercially confidential, indigenous, security-restricted), the legal basis for processing, and the technical and organisational measures applied.
Sensitive-data repository
A repository specifically designed to hold sensitive research data — typically personal data, health data, criminal-justice data, commercially-confidential data, or culturally-sensitive Indigenous data — with enhanced access controls, audit logging, contractual access conditions, and (often) a secure analysis environment.
Trusted research environment
A secure computing environment — typically delivered as a remote-access workspace with controlled inbound/outbound data flows — that allows accredited researchers to analyse sensitive data in situ without exporting the data, supporting privacy-preserving secondary research use.

Metadata and exchange standards

The schemas, protocols, and serialisations that let research-information systems exchange records without manual mapping.

For the applied guidance behind these terms, see metadata, provenance and description standards.

20 terms

CERIF
Common European Research Information Format: an EU-recommended data model and exchange schema for research information, developed and maintained by euroCRIS, that defines core entities (Person, Project, Publication, OrgUnit, Funding, Equipment) and the relationships among them with explicit role-and-time semantics.
Crossref deposit XML
The XML schema family maintained by Crossref that publishers and content-registration members use to submit metadata when minting Crossref DOIs, with distinct sub-schemas for journals, books, conference proceedings, preprints, peer reviews, grants, and standards.
DataCite metadata schema
The XML/JSON schema (current version 4.x) maintained by the DataCite Metadata Working Group that defines the mandatory, recommended, and optional metadata properties to be supplied when registering a DOI through DataCite.
DMP machine-readable expression
The structured serialisation of a DMP's content in a format (JSON, JSON-LD, XML) that conforming software can parse, validate, and act upon without human re-interpretation.
DMP RDA-JSON-LD
A JSON-LD context and ontology rendering of the RDA DMP Common Standard, enabling DMP graphs to be expressed as linked data and merged with other research-information graphs.
Identifier crosswalk
A mapping or correspondence table between identifiers in different schemes that refer to the same real-world entity, allowing a system holding one scheme's identifier to find the equivalent identifier in another scheme.
Identifier scheme
A named, formally-defined system for constructing, issuing, and resolving identifiers, typically defined by a syntax, an authority structure (who can mint), a metadata schema, and a resolution policy.
Identifier syntax
The formal grammar that constrains the textual form of identifiers in a given scheme — what character sets, lengths, prefixes, separators, and check digits are allowed — typically expressed as a regular expression or BNF grammar in the scheme's specification.
JATS XML
Journal Article Tag Suite (ANSI/NISO Z39.96): an XML vocabulary for describing the textual content and metadata of scholarly journal articles, comprising tag sets for archival (Green), publishing (Blue), and authoring (Pumpkin) use cases.
Linked Data Fragments
A family of Web interfaces for publishing RDF data — ranging from data dumps through SPARQL endpoints to lightweight Triple Pattern Fragments — designed to allow clients to query large RDF datasets without overloading the server.
METS
Metadata Encoding and Transmission Standard: a Library of Congress-maintained XML schema for encoding descriptive, administrative, and structural metadata for digital library objects, providing a wrapper around content files and metadata sections.
MODS
Metadata Object Description Schema: a Library of Congress XML schema for descriptive bibliographic metadata, designed as a richer alternative to Dublin Core and a more flexible alternative to MARC for library, archive, and repository contexts.
OAI-ORE
Open Archives Initiative Object Reuse and Exchange: a set of specifications, complementary to OAI-PMH, that define how to describe and exchange aggregations of Web resources (a 'compound digital object') using standard Web Architecture and Linked Data principles.
OAI-PMH
Open Archives Initiative Protocol for Metadata Harvesting: a low-barrier HTTP/XML protocol that allows a 'data provider' system (typically a repository or CRIS) to expose its metadata records for incremental harvesting by 'service providers' (aggregators, search services, national portals).
PID metadata
The descriptive metadata registered with a PID provider at the time of identifier minting, and updated thereafter, that describes the identified entity — its title, creators, dates, types, related identifiers, and so on — and that is exposed by the provider's APIs alongside the identifier itself.
RDA DMP Common Standard
A community-developed application profile, maintained by the Research Data Alliance DMP Common Standards Working Group, defining the core entities, properties, and JSON schema for expressing machine-actionable DMPs.
RDF triple store
A database management system specialised for storing and querying RDF (Resource Description Framework) statements — subject-predicate-object triples (or named-graph quads) — typically with a SPARQL query engine on top.
ResourceSync
ANSI/NISO Z39.99: a framework for synchronising resources between a source server and one or more destination servers using Web-based sitemaps with extensions for change lists, incremental updates, and push notifications.
RIOXX
Research Outputs Information Exchange XML (RIOXX): a UK-originated application profile for institutional-repository metadata, layered on Dublin Core, designed to capture funder, project, and licence information for OA-compliance reporting, currently maintained at v3.0.
SPARQL endpoint
An HTTP endpoint that accepts SPARQL queries (the W3C-standard query language for RDF) against an RDF dataset and returns results in standard formats (XML, JSON, CSV, TSV), implementing the SPARQL 1.1 Protocol.

Repositories and infrastructure

Where research data, code, and samples are deposited, preserved, and certified for long-term trust.

For the applied guidance behind these terms, see repository selection, preservation and storage infrastructure.

211 terms

3-2-1 Backup Strategy
A data-protection rule requiring that any dataset worth protecting exist as at least three total copies, stored across at least two different storage media or systems, with at least one of those copies kept at a location or system genuinely separate from where the working copy lives. All three conditions must hold simultaneously -- extra copies on the same medium, or an offsite copy with no independent local redundancy, do not satisfy the rule on their own.
ACM Artifact Review and Badging
ACM Artifact Review and Badging is a formal peer-evaluation process, run alongside conventional paper peer review at ACM-sponsored conferences and in ACM journals, in which authors of a computer-science research paper voluntarily submit their supporting 'artifacts' -- source code, datasets, containers, build scripts, or other digital objects needed to run the reported experiments -- to a separate Artifact Evaluation Committee (AEC). The AEC checks whether the artifact is documented and runnable, and in stricter tracks, whether running it actually reproduces the paper's reported results. A paper that passes earns one or more standardized badge icons (Artifacts Available, Artifacts Evaluated -- Functional, Artifacts Evaluated -- Reusable, Results Reproduced, Results Replicated), printed on the published paper itself as a visible, per-paper certification distinct from the paper's own peer-review outcome.
Administrative Metadata
Administrative metadata is the category of metadata used to manage a resource — as distinct from metadata used to describe or discover it. It typically covers four practical areas: rights and permissions (licensing, access restrictions), provenance and acquisition history (who deposited it, when, and how), technical/file-format details (file type, software/equipment used, checksums), and access-control information. NISO's standard three-way split names it alongside descriptive and structural metadata, and treats technical and preservation metadata as administrative subtypes — though usage varies slightly by standard: METS, for example, models technical, rights, provenance, and preservation information together under a single amdSec (administrative metadata section), sitting alongside — not nested inside — its descriptive metadata section (dmdSec).
Affiliation (CRIS)
In CRIS terminology, a time-bounded relationship between a Person record and an Organization Unit, capturing the role (employee, visitor, honorary, student), start and end dates, employment fraction, and other contextual attributes of the association.
ANSI/NISO Z39.19 (Controlled Vocabularies and Thesauri)
A controlled vocabulary or thesaurus is constructed to ANSI/NISO Z39.19-2005 (R2010), "Guidelines for the Construction, Format, and Management of Monolingual Controlled Vocabularies," when it follows the standard's rules for term selection and form (preferring noun phrases, singular/plural conventions, avoiding ambiguous homographs without a qualifier), and when it displays the three relationship types Z39.19 defines between terms: equivalence relationships (USE/UF, linking a non-preferred synonym to its preferred term), hierarchical relationships (BT/NT, Broader Term/Narrower Term, for genus-species, whole-part, or instance relationships), and associative relationships (RT, Related Term, for terms that are conceptually related but not synonymous or hierarchical). The standard covers several vocabulary types along a spectrum of structural complexity -- authority lists, synonym rings, taxonomies, and full thesauri -- and specifies how each should be tested, formatted, and maintained over time (adding, deprecating, and merging terms as usage and the underlying domain evolve).
ANSI/NISO Z39.56 (SICI — Serial Item and Contribution Identifier)
A code counts as a SICI only if it assembles all three ANSI/NISO Z39.56-1996 [R2002] segments in sequence — an Item segment (ISSN + chronology + enumeration), a Contribution segment (starting page + title-code checksum for the specific article), and a Control segment (structure/version/check-character) — to identify one specific item within a serial. NISO withdrew the standard in 2012; new article-level identification now runs through the DOI.
ARC Research Management System (RMS)
The Research Management System (RMS) is the Australian Research Council's web-based platform for administering the full lifecycle of applications to its National Competitive Grants Program (NCGP): preparing and submitting applications, assigning and recording assessor evaluations, handling applicant rejoinders, announcing outcomes, and managing the resulting grant agreement, including post-award activity such as variations, annual/end-of-year reporting, and final reports. It is accessed at rms.arc.gov.au by registered users, and Administering Organisations (AOs) -- the Australian universities and other eligible institutions through which ARC grants are held -- are generally required to submit NCGP applications through RMS unless the ARC advises otherwise for a specific scheme.
ARRIVE 2.0
The 2020 revision of the Animal Research: Reporting of In Vivo Experiments guidelines, a checklist of items that should be reported in any publication describing animal research, in order to enable assessment and replication.
Australian Research Data Commons (ARDC)
The national research data infrastructure provider for Australia: a not-for-profit company, funded through the Australian Government's National Collaborative Research Infrastructure Strategy (NCRIS), that builds, operates, and stewards shared digital research infrastructure — including platforms, skills programs, and persistent-identifier services such as RAiD — for use across the Australian university and public-research sector rather than by any single institution.
Authentication of key resources
The verification, by methods appropriate to the resource type, of the identity and integrity of biological and chemical materials used in research, including cell lines, antibodies, animal models, and specialty chemicals.
Biobank
An organised collection of biological samples (typically human samples such as blood, tissue, DNA, urine) together with their associated clinical, demographic, and lifestyle data, governed for use in biomedical research.
Biorepository
A facility or organisation that collects, processes, stores, and distributes biological materials and their associated data for research, encompassing both human and non-human samples, distinguished from a 'biobank' by usage in some communities to denote broader scope or specific research projects.
Blinding and Masking in Clinical Trials
Blinding (or masking, the term FDA/ICH guidance now prefer) in a clinical trial is the procedure of keeping one or more trial parties -- the participant, the investigator/care provider, the outcomes assessor, and/or the data analyst -- unaware of which treatment arm a given participant was assigned to, from the point of randomization onward. It is recorded on ClinicalTrials.gov as the "Masking" data element with levels of None (open-label), Single, Double, Triple, or Quadruple, depending on how many of those roles are kept unaware. Masking is a distinct methodological safeguard from randomization and allocation concealment: it addresses performance bias (differential behavior or care from knowing the assignment) and detection/ascertainment bias (differential outcome judgment), not selection bias at enrollment.
Bring Your Own Device (BYOD) for Research Data Collection
In research data collection, BYOD (bring your own device) is a practice in which participants or researchers use their own personally-owned smartphone, tablet, or wearable -- rather than a device supplied, configured, and owned by the study -- to run a data-collection app or sensor. What makes an instance BYOD is device ownership and control: the hardware belongs to the participant or researcher, not the study team, so the study cannot fully standardize, lock down, or reclaim it, and must build its data-integrity and security controls around that constraint.
ClinicalTrials.gov
ClinicalTrials.gov is the NLM-operated (NIH, with FDA) public registry and results database for clinical studies, distinguished from clinical trial registration generally by being the specific U.S. system launched in 2000 (FDAMA Sec. 113) whose mandatory-registration and results-reporting requirements were established by the FDAAA 801 final rule (effective 18 January 2017), and which assigns every registered study a unique NCT identifier.
Cochrane Handbook for Systematic Reviews of Interventions
The Cochrane Handbook for Systematic Reviews of Interventions is the methodological standard maintained by Cochrane (founded 1993, named after epidemiologist Archie Cochrane) that specifies how a systematic review of a health intervention must be planned, conducted, and reported to qualify as a Cochrane Review. It requires a registered protocol with a PICO-structured question, a comprehensive pre-specified search strategy across multiple databases, dual independent screening and data extraction, formal risk-of-bias assessment (the RoB 2 tool for randomized trials, ROBINS-I for non-randomized studies), GRADE-based certainty-of-evidence rating, and, where studies are sufficiently similar, meta-analysis. Reviews that follow the Handbook are published in the Cochrane Database of Systematic Reviews (CDSR) within the Cochrane Library and are maintained as living documents, updated as new evidence emerges.
Code repository
A version-controlled storage location for source code, typically operated on top of a distributed version-control system such as Git, exposing the code's full revision history, branches, tags, and (often) collaboration features such as issues, pull requests, and code review.
Code review (research software)
A structured review of research software by one or more peers, focused on correctness, clarity, documentation, testing, and fitness for purpose, conducted before publication or as part of community-curated software repositories.
Concordat on Open Research Data
The Concordat on Open Research Data is a UK sector-wide policy framework, published 28 July 2016 by the Higher Education Funding Council for England (HEFCE), Research Councils UK (RCUK, since absorbed into UK Research and Innovation, UKRI), Universities UK, and the Wellcome Trust, setting out ten principles for how researchers, research organisations, and funders should approach making research data openly available. It is not a regulation, mandate, or legal instrument: an institution, funder, or research group operationalises the Concordat when its research data policy, Data Management Plan (DMP) guidance, or grant terms explicitly build in the ten principles — for example, requiring a documented, case-by-case justification for any restriction on data openness (Principle 2), recognising a researcher's right to a reasonable first-use period before data must be shared (Principle 4), or expecting data underlying a publication to be accessible, citable via a persistent identifier, and retained for at least ten years from the publication date (Principle 8).
Confidence Interval
A confidence interval is a range of values, calculated from sample data using a defined estimation procedure (point estimate plus a margin of error derived from the sampling distribution of the estimator), constructed at a stated confidence level (commonly 95%). The confidence level describes the long-run reliability of the construction procedure across repeated sampling -- the proportion of such intervals that would contain the true, fixed population parameter if the sampling were repeated many times -- not the probability that this specific interval contains the parameter. A range lacking a stated confidence level and a defined estimation procedure (e.g. a bare sample min-max) is not a confidence interval.
CONSORT 2010
The 2010 edition of the Consolidated Standards of Reporting Trials, a 25-item checklist and participant-flow diagram covering items that should be reported in any randomised controlled trial publication.
Converis (Clarivate)
Converis is Clarivate's commercial CRIS (Current Research Information System) product: a configurable platform that ingests publication and bibliometric data (including from Web of Science and other sources), manages pre-award and post-award research-project workflows, and exposes a public research-portal layer for institutions, funders, and national research agencies. An institution's software counts as an instance of Converis specifically when it is the Clarivate-branded, vendor-hosted or vendor-configured platform -- not merely any locally-built research-activity database with similar functions.
Core metadata
The minimal, standard set of descriptive fields — typically title, creator, date, persistent identifier, subject/keywords, format, and rights — that a dataset or research output must carry in order to be findable and citable, independent of any richer, discipline-specific metadata layered on top of it.
CoreTrustSeal
A community-based, non-profit certification scheme for trustworthy data repositories, operated by the CoreTrustSeal Foundation, awarded against 16 published requirements covering organisational infrastructure, digital object management, and technical infrastructure.
COUNTER 5
COUNTER 5 (formally the COUNTER Code of Practice, Release 5, or 'COP5') is the current international standard governing how publishers, aggregators, and other content platforms must record and report usage statistics — searches, investigations, and requests — for the electronic resources (journals, databases, books, and multimedia items) they license to libraries and consortia. A usage report qualifies as 'COUNTER 5 compliant' only if it is produced from one of COP5's defined Master Report templates, uses COP5's controlled Metric_Types vocabulary, applies COP5's double-click and robot-filtering deduplication rules consistently, and (where automated) is retrievable via the COUNTER API (formerly called SUSHI, the Standardized Usage Statistics Harvesting Initiative). The standard exists so that a usage number from one publisher platform means the same thing as the equivalent number from another, letting libraries compare cost-per-use and make collection-development and renewal decisions on a like-for-like basis.
CRIS
Current Research Information System: an enterprise-class software system that aggregates, stores, and publishes information about a research organisation's activities — its researchers, publications, projects, funding, equipment, collaborations, and outputs — and exposes that information to internal management and external reporting consumers.
CRIS interoperability
The capacity of CRIS systems to exchange data with each other and with adjacent systems (repositories, funders, publishers, aggregators) through shared data models, schemas, protocols, and persistent identifiers — most prominently CERIF, OAI-PMH, OpenAIRE Guidelines, and PID-based joins.
CrossMark
A Crossref member service in which a publisher embeds a button on a published item (its HTML landing page, PDF, and/or ePub) that a reader can click to check the item's current status directly from within the item itself, rather than having to separately search the publisher's site. Clicking the button opens a panel showing whether the record has been updated since original publication -- including a correction, retraction, expression of concern, or one of the other editorially significant update types Crossref's schema defines -- alongside supporting metadata such as key dates (submission, revision, acceptance), funder information, licence terms, handling editors, and peer-review information where the publisher has chosen to provide it. A publication only counts as CrossMark-enabled if its publisher has (a) registered a DOI-assigned update policy, (b) referenced that policy in the item's Crossref metadata record even when no update yet exists, and (c) implemented the button so the status check happens at the item, not just on a separate corrections page a reader would have to know to look for.
Crossref Event Data
Crossref Event Data was a free Crossref service that captured and made available structured records of "events" -- online mentions such as tweets, Wikipedia citations, newspaper articles, blog posts, and Wikipedia edits -- that referenced scholarly content identified by a DOI. Each event linked a source (e.g. a specific tweet or Wikipedia page) to a subject (a DOI-identified research output) with a timestamp, and the resulting data fed into third-party altmetrics tools and dashboards. Crossref formally sunset the service: the public Event Data API and related infrastructure stopped being available from 23 April 2026, after roughly a decade in operation, and it should now be described in the past tense as a discontinued service, not an active one.
Crossref Simple Text Query
<p>Crossref Simple Text Query is a free web-based matching tool at <a href='https://doi.crossref.org/simpleTextQuery' target='_blank' rel='noopener'>doi.crossref.org/simpleTextQuery</a> that takes an unstructured reference or reference list &#8212; pasted as plain citation text, not structured fields &#8212; and returns the matching DOI for each item it can identify. A submission can hold up to 1,000 references, entered one per line or block, and the tool attempts to return exactly one best-match DOI per reference by default (an option to list all possible matches is available when a reference is ambiguous). No Crossref membership or account is required to use it, and there is no charge. The same underlying free-text matching is also exposed programmatically through the Crossref REST API's <code>query.bibliographic</code> parameter on the <code>/works</code> route (for example <code>api.crossref.org/works?query.bibliographic=your+citation+string</code>), which is the machine-actionable equivalent of pasting a citation into the web form &#8212; useful for batch reference-linking pipelines rather than one-off lookups.</p>
Crowdsourced replication
A coordinated effort in which many independent laboratories or teams attempt to replicate the same set of studies under pre-specified protocols, in order to estimate field-wide replicability.
Darwin Core
A data standard ratified by TDWG (Biodiversity Information Standards) — a glossary of defined terms organized into a small set of classes (Occurrence, Taxon, Event, Location, MaterialEntity, Identification, and others) — used to structure and exchange biodiversity data documenting the occurrence of organisms and the specimens or observations that record them. A dataset counts as Darwin Core data when its fields are actually mapped to defined Darwin Core terms (e.g. dwc:scientificName, dwc:eventDate, dwc:decimalLatitude), not merely because it describes species or specimens in some other format.
Data Curation
The active, ongoing management of research data across its lifecycle -- organizing files into a coherent structure, describing them with standardized metadata, validating and cleaning values, and taking preservation actions such as format migration and fixity checking -- carried out specifically so the data remains findable, interpretable, and reusable by someone other than its original creator, long after the project that produced it ends. Curation is a defined set of actions applied to data over time; it is distinct from simply storing a copy of it.
Data Curation Network (DCN)
The Data Curation Network (DCN) is a membership consortium of academic and non-profit research-data repositories that pools trained data curators across member institutions so that any participating library can call on specialist review for a dataset, rather than each institution having to hire full curation expertise for every data type it receives. It is distinct from 'data curation' the generic activity and 'data curator' the generic role: DCN is a specific, named organization, hosted administratively by the University of Minnesota Libraries, that operationalizes those generic concepts as a shared-staffing service across its member institutions.
Data Curator
A data curator is an institutional role — usually in a library, research data service, or research-computing unit — that prepares research data for deposit, applies metadata standards, checks FAIR readiness, and advises researchers on repository selection and data management plans, distinct from owning the underlying research itself.
Data Dredging
Searching a dataset for statistically significant relationships across many variables, subgroups, or model specifications without a hypothesis specified in advance, then reporting whichever pattern crosses a significance threshold as if it had been the study’s pre-specified object of investigation, without correcting for the number of comparisons actually performed.
Data Ownership
The allocation of legal title to and decision-making control over a research dataset. Because raw data is rarely copyrightable and no single U.S. statute assigns default ownership of research data, the operative answer in practice comes from a combination of institutional policy, the funder's terms and conditions, the Data Management Plan, and any executed data sharing/use agreement -- not from a single ownership doctrine. Distinguish from data stewardship/custodianship, which describes accountability for a dataset's day-to-day management regardless of who holds title.
Data publication platform
A platform that supports the publication of research data as a citable artefact — assigning a persistent identifier, presenting a landing page, and applying review, curation, or peer-review processes — distinct from purely depositional storage.
Data Repository
A system or platform dedicated to the long-term storage, curation, preservation, and dissemination of research data, providing persistent identifiers, standardized metadata, and defined access/reuse terms -- distinct from general-purpose file storage (a network drive or cloud folder), which does not guarantee a dataset stays findable, citable, or usable once active work on it ends.
Data Stewardship Wizard (DSW)
An open-source, self-hostable DMP-authoring platform that generates data management plans from versioned, branching questionnaire 'knowledge models' rather than static funder templates, developed and stewarded as an ELIXIR infrastructure tool and released under the Apache License 2.0.
Dataset landing page
The human-readable web page that a dataset's persistent identifier (typically a DataCite DOI) resolves to, presenting the dataset's title, creators, description, identifiers, dates, version history, related works, access conditions, and a link to download or request the data.
dbGaP (Database of Genotypes and Phenotypes)
The NIH/NCBI repository for genotype-phenotype study data (GWAS results, sequencing/omics data, linked phenotype data), split into an open-access tier (study documentation, summary statistics) and a controlled-access tier (de-identified individual-level genotype and phenotype records) gated by Data Access Committee review of a Data Access Request -- distinct from a general sequence archive because the individual-level linkage, not just the data type, is what triggers controlled access.
DDI Codebook
A metadata document is an instance of DDI Codebook (DDI-C) when it is a well-formed XML file rooted at the <codeBook> element and containing the DDI Alliance's standard sections for describing a single dataset: a study description (title, authors, abstract, methodology, sampling procedure), a file description of the physical data file, and a data description that documents each variable individually, including its name, label, question text, value labels, and universe. It is maintained by the DDI Alliance and versioned (current releases in the 2.x line, e.g. 2.5/2.6); it is the single-dataset-focused branch of DDI, distinct from the multi-wave/lifecycle-tracking DDI Lifecycle (DDI 3.x) specification.
Decentralized Clinical Trial (DCT) Platform
A decentralized clinical trial (DCT) platform is the software and connected-service layer that lets some or all trial activities happen away from a traditional investigative site, in line with FDA's trial-conduct model for decentralized elements. In practice a DCT platform is not one thing but an integration of several components under a single participant- and site-facing interface: telehealth/video-visit tooling for remote investigator encounters, eConsent for remote informed consent, home health or mobile-nursing coordination for in-home visits and specimen collection, connections to local or community labs and pharmacies as alternate data-collection or dispensing sites, direct-to-patient (DtP) investigational product (IMP) shipment and chain-of-custody tracking, and eCOA/ePRO or connected devices for participant-reported and sensor-derived data — usually feeding a central EDC and safety database in real time. A trial "uses a DCT platform" when a vendor's connected toolset, rather than a single point solution, is the operational backbone for how remote or hybrid visits are scheduled, conducted, documented, and reconciled against the protocol. The defining administrative complication is that decentralizing activities does not decentralize accountability: the sponsor and the trial's IRB(s) of record remain responsible for GCP compliance, and each local nurse, telehealth clinician, lab, or pharmacy performing a delegated trial activity must be added to the delegation-of-authority log, must practice within a jurisdiction where they are licensed, and must be covered by the same monitoring and oversight expectations as staff at a traditional site.
Dimensions (Research Database)
A record is part of Dimensions when Digital Science's Dimensions platform has ingested it — as a publication, grant, patent, clinical trial, dataset, or policy document — through its largely automated ingestion pipeline (funder data feeds, publisher metadata, patent-office records, clinical-trial registries, and full-text/acknowledgment mining), rather than through the title-by-title editorial selection Scopus and Web of Science apply. Where a funding relationship between a grant and the outputs it produced can be identified, Dimensions represents that link directly in its data model — this explicit funding-to-output linkage is the feature that most consistently distinguishes a Dimensions record from an equivalent Scopus or Web of Science record.
Discipline-specific repository
A repository whose scope is bounded to a particular research discipline or sub-discipline, with curation practices, metadata schemas, and community standards tailored to that domain's data types, terminologies, and norms.
DMPTool
DMPTool is a free, web-based data management plan (DMP) authoring service operated by the California Digital Library (CDL), through its University of California Curation Center (UC3), on behalf of an international partner consortium. It runs on the open-source DMPRoadmap codebase -- the same platform underlying the Digital Curation Centre's UK-based DMPonline -- and is the DMPRoadmap deployment used predominantly by US institutions and funders. An instance qualifies as DMPTool (rather than a generic DMPRoadmap deployment or DMPonline) when it is accessed at dmptool.org, uses CDL/UC3-maintained funder templates (NSF, NIH, NEH, DOE, IMLS, and others), and is administered under an institution's DMPTool participation agreement rather than a DCC/DMPonline subscription.
DOAJ (Directory of Open Access Journals)
DOAJ (Directory of Open Access Journals) is an independent, community-curated online directory that indexes peer-reviewed open-access journals meeting its published inclusion criteria. A journal counts as "DOAJ-indexed" only once it has applied and DOAJ's editorial team has evaluated and accepted that application against DOAJ's own checkable standards for genuine open access, documented peer review, an identifiable editorial board, ISSN registration, licensing/copyright transparency, and public disclosure of any author-facing fees -- self-describing as open access, or being under DOAJ review, is not the same claim as being listed.
DOAJ Native XML schema
DOAJ Native XML is the XML metadata format that DOAJ (Directory of Open Access Journals) defines and maintains for publishers to submit article-level metadata directly into DOAJ's index: title, authors and their ORCID iDs, affiliations, abstract, DOI, ISSN(s), volume/issue/page, publication date, and full-text URL. A file counts as an instance of this schema when it validates against DOAJ's own XSD (not a third-party schema) and arrives through one of DOAJ's inbound submission channels — file upload via a publisher's DOAJ account, DOAJ's API, or an OJS (Open Journal Systems) export plugin that generates DOAJ-conformant XML automatically. It is one of two XML dialects DOAJ accepts for article submission, the other being Crossref XML (versions 4.4.2 or 5.3.1); the two cannot be mixed within a single upload file, and a Crossref XML submission must carry an ISSN for DOAJ to process it.
Domain repository
Synonym for discipline-specific repository: a repository whose scope is a particular research domain (or domain-sub-area), with curation practices and metadata tailored to that domain.
Dryad (concept)
A non-profit generalist research data repository operated by Dryad Data Inc. (in partnership with the California Digital Library) that publishes peer-reviewed-paper-linked datasets, mints DataCite DOIs, and applies curation review before publication.
DSpace-CRIS
An open-source extension of the DSpace repository platform, developed by 4Science and the DSpace community, that adds CRIS-style entity management for researchers, projects, organisational units, journals, and other research-information entities alongside the existing repository content.
Dublin Core
Dublin Core is the cross-domain metadata vocabulary, governed by the Dublin Core Metadata Initiative (DCMI), built around fifteen core resource-description elements (Title, Creator, Subject, Description, Publisher, Contributor, Date, Type, Format, Identifier, Source, Language, Relation, Coverage, Rights) formally standardized as ISO 15836, ANSI/NISO Z39.85, and IETF RFC 5013 (Simple Dublin Core), plus the larger DCMI Metadata Terms ("dcterms:") vocabulary that adds element refinements, encoding schemes, and additional properties/classes for more precise description (Qualified Dublin Core). A metadata record is a Dublin Core record when its fields map to this DCMI-maintained element set or the extended dcterms vocabulary, rather than to an unrelated, locally invented field set or an unrelated domain-specific schema.
Ecological Metadata Language (EML)
The XML-based metadata specification (current stable release EML 2.2.0) used to document a research dataset, or an associated literature/software/protocol resource, as a single machine-readable package: its identity and citation information, collection methods and sampling design, taxonomic/geographic/temporal coverage, and the structure and attributes of its constituent data tables or other entities. Something is an EML record if it validates against the EML XML schema maintained via the NCEAS/eml specification and describes a dataset (or related resource) at the package level, rather than encoding the individual records within that dataset.
Effect Size
Effect size is a standardized, quantitative measure of the magnitude of a difference, relationship, or association observed in a study, reported independently of sample size and distinct from statistical significance. Where a p-value indicates only whether an observed effect is unlikely to be due to chance, an effect size (e.g., Cohen's d, Pearson's r, an odds ratio) indicates how large or practically meaningful that effect actually is. The APA Publication Manual (7th edition, Section 6.5) requires reporting an effect size and confidence interval for each primary outcome, and most major journals across the behavioral, social, and biomedical sciences hold manuscripts to an equivalent expectation.
Electronic Data Capture (EDC)
Electronic Data Capture (EDC) is the general category of software systems that collect clinical or research study data directly in electronic form -- via electronic case report forms (eCRFs) -- at the point of collection, in place of paper-based data collection. A system counts as EDC when it provides study-specific eCRFs, direct point-of-collection data entry, field-level validation and query management, and an audit trail sufficient to make the electronic record the study's data of record. EDC is the site-facing data-entry layer within the broader discipline of clinical data management; a full Clinical Data Management System (CDMS) is the broader term covering EDC plus surrounding query-management, coding, and database-lock workflow. EDC is distinct from a Clinical Trial Management System (CTMS), which manages trial operations rather than the research data itself. Both non-profit/academic platforms (e.g., REDCap) and commercial, enterprise-scale platforms (e.g., Medidata Rave EDC, Oracle Clinical One) fall within the EDC category.
Electronic Lab Notebook (ELN)
An electronic lab notebook (ELN) is software that replaces the paper laboratory notebook as the primary record of experimental work. An entry qualifies as ELN record-keeping when it is timestamped and attributed at the moment of creation, when edits are versioned rather than overwritten (so a full history remains retrievable), and when the record is structured enough to be searched, exported in a non-proprietary format, and -- increasingly -- assigned a DOI and deposited as a citable research output.
Electronic Medical Record (EMR)
An electronic medical record (EMR) is the digital equivalent of a single healthcare provider's paper chart: a patient's medical and treatment history as captured, stored, and used within one clinical practice or healthcare organization. An artifact is an EMR (rather than an EHR) when its data model, access controls, and interoperability are scoped to a single organization's internal clinical workflow, not designed for structured exchange with external providers, payers, or research systems -- even if the underlying software vendor also sells EHR-branded products elsewhere.
Electronic Theses and Dissertations (ETD)
A thesis or dissertation prepared, submitted, and archived in digital form (typically a PDF plus any supplementary files) rather than only as a bound print copy, submitted to satisfy a degree requirement. A work qualifies as an ETD once digital deposit — most often to an institutional repository, and often onward to a national or international aggregator such as NDLTD — is the accepted or required submission format, even if the institution also retains a print copy alongside it.
Elements (Symplectic)
A commercial CRIS product developed by Symplectic (part of Digital Science) that automates the discovery, capture, and management of research outputs and activities for individual researchers and institutional reporting workflows, with strong emphasis on researcher-facing workflows.
Elicit (AI Research Assistant)
Elicit is a named AI research-assistant platform (built by the public benefit corporation Elicit, spun out of the nonprofit lab Ought in 2023) that searches an indexed academic-literature corpus and performs structured evidence-extraction tasks -- summarizing papers, extracting and tabulating data points across many papers with sentence-level source citations, and supporting systematic-review-style screening and data-extraction workflows aligned to PRISMA 2020. It is a specific product, not a generic label for "AI that helps with research" -- distinct from citation-graph discovery tools (Connected Papers, ResearchRabbit) and general-purpose AI writing assistants -- and its extraction/screening output requires independent verification against the source papers rather than being treated as ground truth.
EOSC
European Open Science Cloud: an EU-led initiative and emerging federation of research data infrastructures intended to provide European researchers with seamless, cross-border, cross-discipline access to data, services, and computational resources under FAIR and open-science principles.
EOSC Federation
The architectural model adopted by EOSC for federating heterogeneous national, thematic, and pan-European research-data infrastructures into a single user-facing layer, with shared identity and access management, monitoring, accounting, helpdesk, and service onboarding.
EOSC Marketplace
The catalogue of FAIR research services accessible to European researchers through the European Open Science Cloud, where service providers register their offerings and researchers can discover, order, and (where applicable) access services with EOSC-federated authentication.
Esploro (Ex Libris)
Esploro is a commercial current research information management (RIM/CRIS) platform sold by Ex Libris (a Clarivate company). Institutions use it to aggregate and showcase research assets -- publications, datasets, and other outputs -- captured from sources such as ORCID, Scopus, Web of Science, PubMed, and Crossref, alongside researcher profiles, and to expose that data through public-facing research portals, funder/compliance reporting, and open-access workflows. It can be deployed standalone or alongside Ex Libris's Alma library services platform, though the vendor markets it as usable independently of Alma.
euroCRIS
euroCRIS is the international not-for-profit membership association of research-information professionals and institutions that holds custodianship of the CERIF data model and coordinates community infrastructure (the DRIS directory, the CRIS Conference, membership meetings, working groups) for interoperability among Current Research Information Systems. Something qualifies as "euroCRIS" only if it refers to the organisation itself — its governance, membership, task groups, and services — not the CERIF standard or a CRIS software product it helps make interoperable.
Extended level certification (trustworthy repository)
The middle tier of a three-level trust framework for research-data repositories, positioned above core level (CoreTrustSeal, a peer-reviewed self-assessment against 16 requirements) and below formal level (ISO 16363, a full external audit). A repository holds extended level certification when it has been awarded the nestor Seal for Trustworthy Digital Archives — a plausibility-checked self-assessment against the 34 criteria of the German standard DIN 31644, administered by the nestor competence network at the Deutsche Nationalbibliothek. It is more rigorous than a core-level self-assessment (more criteria, external plausibility review of the evidence submitted) but, unlike formal level, does not involve an accredited external auditor conducting an on-site or fully independent audit.
FAIR Data Principles
The FAIR data principles are a set of four foundational goals — Findable, Accessible, Interoperable, and Reusable — for how research data and its metadata should be described, deposited, and structured so that both humans and machines can locate, retrieve, and reuse it with minimal manual effort. FAIR is a set of guiding goals for data stewardship, not a certification, a checklist with one correct implementation, or a data-sharing mandate in itself: a dataset is 'FAIR' to the degree its metadata and deposit environment satisfy the fifteen sub-principles across the four categories, and different repositories, disciplines, and data types satisfy them by different concrete means.
FAIRsharing (concept)
A curated, community-driven registry of databases, standards (metadata, identifiers, formats, terminologies), and data policies relevant to research data, maintained at the University of Oxford with linkage to funders, journals, and standards organisations.
Figshare (concept)
A commercial generalist research repository operated by Digital Science that accepts datasets, figures, presentations, papers, software, and other research artefacts, minting DataCite DOIs and offering institutional-branded instances ('Figshare for Institutions') alongside the public service.
Forking paths
The phenomenon by which the cumulative effect of many small, data-contingent analytical choices inflates false-positive rates even when each individual choice appears defensible.
Funding entity
In CRIS terminology, an entity representing a specific award or funding instance — its funder, award number, amount, currency, start and end dates, and the project and people it supports — distinct from the abstract Funder organisation and from the project itself.
GA4GH (Global Alliance for Genomics and Health)
An international not-for-profit alliance, founded in 2013, that develops and maintains technical interoperability specifications and policy frameworks (such as the Data Use Ontology, GA4GH Passport, Beacon, htsget, Crypt4GH, and Phenopackets) enabling genomic and health-related data to be shared responsibly across research and clinical institutions worldwide; GA4GH does not host data itself but defines the standards that repositories, biobanks, and national genomic initiatives implement to interoperate.
Garden of forking paths
The Gelman-Loken metaphor for the implicit, data-dependent multiplicity of analytical choices made in the course of an empirical study, even by analysts not engaged in explicit p-hacking.
Gateway to Research (GtR)
UKRI's free, public web portal (gtr.ukri.org) that publishes structured data on UKRI-funded research and innovation projects -- overview, organisations, people, publications, and outcomes -- for all seven UKRI research councils and Innovate UK, distinct from Researchfish, the grantee-facing system through which much of that outcome data is originally submitted.
GBIF (Global Biodiversity Information Facility)
GBIF (the Global Biodiversity Information Facility) is an intergovernmental network and open-access data infrastructure -- coordinated by a Secretariat in Copenhagen and funded by its participating governments and organizations -- that aggregates, indexes, and republishes species-occurrence and specimen records contributed by data-holding institutions worldwide through a global network of national and thematic Participant Nodes. A dataset counts as GBIF-mediated when it has been registered through a GBIF Participant Node (or directly via GBIF.org) and is discoverable, searchable, and downloadable through the GBIF occurrence-search interface and API; the great majority of datasets reach GBIF as Darwin Core Archives (DwC-A), most commonly built and published using GBIF's own Integrated Publishing Toolkit (IPT). GBIF itself does not define a data standard -- it consumes Darwin Core, the pre-existing TDWG biodiversity-data standard -- and is best understood as the largest aggregating platform built on top of that standard, not the standard itself.
Generalisability
The extent to which a study's findings extend to populations, settings, or conditions other than those directly sampled.
Generalist repository
A repository that accepts research outputs from any discipline, applying domain-agnostic curation and discovery, and serving as a deposit destination for outputs that have no natural discipline-specific home or whose authors prefer a single multidisciplinary venue.
GitHub mirror
A copy of a Git repository (or set of repositories) hosted on GitHub that tracks an upstream source repository elsewhere, typically maintained for redundancy, visibility, or community-engagement reasons rather than as the canonical primary copy.
Globus
A non-profit, hosted research data management service - built on GridFTP-based transfer technology and operated by the University of Chicago with Argonne National Laboratory - that provides high-performance, fault-tolerant file transfer, institutional 'endpoints' (via Globus Connect Server/Personal), and controlled sharing between research storage systems, HPC centers, and repositories.
Google Scholar API
Google, as a matter of long-standing, publicly stated policy, does not operate an official, sanctioned application programming interface for Google Scholar -- there is no documented endpoint, no API key registration process, and no rate-limited free or paid tier comparable to those offered by Crossref, ORCID, Scopus, or OpenAlex. "Google Scholar API" as a search term almost always refers to one of two distinct things instead: (1) a third-party commercial scraping service (e.g. SerpApi, Apify, Oxylabs) that queries Google Scholar's public web pages on a customer's behalf and returns the result as structured JSON, or (2) an open-source scraping library (e.g. the Python package "scholarly") that automates a browser or HTTP session against the same public pages with no official sanction from Google. Both categories work by parsing Google Scholar's HTML search-results and profile pages, not by calling an endpoint Google built or documents for that purpose, and both are constrained by Google Scholar's robots.txt disallow rules and general Google Terms of Service language prohibiting automated querying without prior permission.
Google Scholar Citations profile
A self-created, self-curated author page on scholar.google.com that displays a researcher's publications alongside auto-updating citation count, h-index, and i10-index figures computed from Google Scholar's own web index, distinct from database-issued author records such as a Scopus Author ID or an ORCID iD.
HARKing (Hypothesising After Results are Known)
Presenting a post-hoc hypothesis, formulated after data analysis, as if it had been the a priori hypothesis under test.
Harvard Dataverse (concept)
A free research-data repository operated by Harvard University on the open-source Dataverse software platform, accepting datasets from researchers worldwide, minting DataCite DOIs, and serving as the flagship instance of the global Dataverse network.
ICPSR (concept)
Inter-university Consortium for Political and Social Research: a consortium-membership-funded data archive based at the University of Michigan that holds and curates over 10,000 social-science research datasets, providing access to member institutions worldwide.
InCites
A subscription research-evaluation and benchmarking platform, owned and sold by Clarivate, built on Web of Science Core Collection publication and citation data, that lets research institutions, funders, publishers and government bodies benchmark research performance, track collaboration, and analyse impact at the level of a researcher, department, institution, country, journal or custom peer group. It is a reporting and benchmarking layer on top of Web of Science rather than a separate bibliographic database of its own -- everything InCites shows is ultimately traceable back to a Web of Science record. It is Clarivate's direct counterpart to Elsevier's Scopus-based SciVal.
Institutional repository
An online, digital collection of research outputs (see Repository) that are connected by their affiliation with a specific institution. Institutional repositories are most commonly associated with universities and other academic organisations, and so the contents of a single institutional repository may therefore cover a range of disciplines. An institutional repository may often be managed as part of a wider suite of services supporting scholarly communication, Open Access and Open Education.
Institutional webpage
A webpage that is associated with the institution at which the author is employed.
Intent-to-Treat (ITT) vs. Per-Protocol Analysis
An analysis-population choice for a randomized trial: an intent-to-treat (ITT) analysis includes every participant in the treatment group to which they were originally randomized, regardless of whether they received, adhered to, or completed that assigned treatment. A per-protocol (PP) analysis instead restricts the analysis to the subset of participants who received the assigned treatment as specified in the protocol, with no major deviations. ITT answers "what happens if this treatment is assigned"; per-protocol answers "what happens if this treatment is actually taken as directed" — and which of the two gives the more conservative estimate depends on whether the trial is testing superiority or non-inferiority.
ISO 16363 (Formal-Level Certification)
The Formal level of the three-tier trustworthy-repository certification framework: a repository is ISO 16363-certified only when an accredited external auditor (typically PTAB, accredited under ISO 17021/16919) has conducted a full audit against ISO 16363's evidence-based metrics and formally certified the result — distinct from CoreTrustSeal's and the nestor Seal's self-assessment models. Full title: Space data and information transfer systems — Audit and certification of trustworthy digital repositories. Descends from the 2007 TRAC (Trustworthy Repositories Audit & Certification) checklist published by RLG/NARA, formalised by CCSDS as ISO 16363:2012 and subsequently revised.
ISO 18626 (Interlibrary Loan Transactions — Application Protocol)
<p><strong>ISO 18626</strong>, <em>Information and documentation — Interlibrary loan transactions</em>, is the international application-protocol standard for the messages that library systems exchange to request, supply, and track resource-sharing loans and copies between institutions. A transaction qualifies as ISO 18626-compliant when it is carried out as a defined sequence of XML messages — a <strong>Request</strong> from the borrowing (requesting) library, and <strong>Supplying Agency</strong> and <strong>Requesting Agency</strong> status/action messages exchanged between the two systems — rather than as free-text email, fax, or a proprietary, non-standard API. First published by ISO in 2014, revised in 2017, and further revised as ISO 18626:2021, the standard is maintained under ISO/TC 46, the technical committee for information and documentation.</p>
ISO 23081 (Records Management Processes — Metadata for Records)
ISO 23081 (Information and documentation — Records management processes — Metadata for records) is the ISO standard that defines the principles and framework governing the metadata records systems must capture so that records can be managed, retrieved, and trusted as evidence throughout their lifecycle. A metadata scheme qualifies as ISO 23081-aligned when it goes beyond simple descriptive metadata (title, creator, date) to capture recordkeeping context: the business process or mandate that created the record, the agents responsible for it, its relationships to other records and aggregations, and its retention/disposition authority. The standard is published in three parts: Part 1 sets out the principles underpinning records metadata; Part 2 gives conceptual and implementation guidance for building a metadata schema consistent with Part 1; Part 3, published as a Technical Report, provides a self-assessment method for evaluating an existing metadata schema against the standard. ISO 23081 was developed by ISO/TC 46/SC 11 as a companion to ISO 15489 (the core international records management standard), translating 15489's principles into a concrete metadata model that records-management systems, archives, and institutional repositories can actually implement.
ISRCTN
<p>ISRCTN &mdash; originally an acronym for International Standard Randomised Controlled Trial Number, now used as the registry's own name &mdash; is a UK-based clinical study registry that assigns a permanent, unique identifier (an ISRCTN number) to a study at or before its start. A study qualifies as "ISRCTN-registered" when its investigator or sponsor has submitted the required minimum data set (design, intervention, population, primary/secondary outcomes, sponsor, funder) through isrctn.com, the entry has passed the registry's validation checks, and it has been assigned a permanent identifier in the format ISRCTN followed by an eight-digit number (for example, ISRCTN12345678). Unlike its original RCT-only scope, ISRCTN now accepts any study assessing the effect of a health intervention on a human population, including non-randomised interventional studies and some observational designs &mdash; the acronym is retained for historical/branding reasons but no longer describes the registry's actual scope.</p>
ISSN
An eight-digit identifier (ISO 3297), formatted NNNN-NNNC with a modulus-11 check digit as the final character, assigned to a continuing resource's TITLE — a journal, magazine, newspaper, or other serial published indefinitely over time, in print or online — by a national ISSN Centre under the coordination of the ISSN International Centre (CIEPS) in Paris. An ISSN identifies the serial as a whole, not any individual article, issue, or volume within it, and each distinct medium of the same title (print vs. online) is assigned its own ISSN, linked together via an ISSN-L.
Je-S (Joint Electronic Submission)
Je-S (Joint Electronic Submission), hosted at je-s.rcuk.ac.uk, is UK Research and Innovation's (UKRI) legacy online system for submitting grant applications and administering research council awards. UKRI's seven research councils used Je-S for close to two decades before progressively replacing it with the UKRI Funding Service (funding-service.ukri.org) from 2023 onward. A form, account, or award is 'in Je-S' when it was created on, or is still being administered through, that older platform rather than the current Funding Service.
Journal Article Reporting Standards (JARS)
Journal Article Reporting Standards (JARS) is the American Psychological Association's set of manuscript-section reporting requirements, published as part of APA Style and incorporated into the <em>Publication Manual of the APA</em> (7th edition). A manuscript is JARS-compliant when the author has followed the module matching its research design: <strong>JARS-Quant</strong> for quantitative studies (numerical data, statistical analysis), <strong>JARS-Qual</strong> for qualitative studies (natural-language, descriptive, or interview-based data), or <strong>JARS-Mixed</strong> for studies combining both. A cross-cutting fourth module, <strong>JARS-Race, Ethnicity, and Culture (JARS-REC)</strong>, adds reporting expectations for how race, ethnicity, and culture are described and analyzed, and applies regardless of which core module a study uses. Each module specifies, section by section (title/abstract, introduction, method, results, discussion), the minimum information a submission must state -- for example, JARS-Quant requires reporting exact sample-size determination and any data exclusions; JARS-Qual requires describing the researcher's own standpoint and role in relation to the data. JARS is APA's own umbrella framework and is distinct from discipline- or design-specific reporting guidelines such as PRISMA (systematic reviews), CONSORT (randomized trials), or COREQ/SRQR (qualitative research specifically) -- a paper can be asked to satisfy JARS alongside one of those other checklists, not instead of it, since JARS covers general manuscript-section content while the others target a specific study design in more procedural depth.
Journal Citation Indicator (JCI)
A journal's Journal Citation Indicator (JCI) is the field-normalized citation-impact score Clarivate calculates for it as part of each annual Journal Citation Reports (JCR) release. It measures the average citation impact of a journal's citable items published in the most recent three-year window, normalized against the average citation rate for all similarly-aged papers of the same document type in the same JCR subject category. A journal qualifies for a JCI if it is indexed in a JCR-eligible Web of Science Core Collection index (SCIE, SSCI, AHCI, or ESCI) with enough citable output in the relevant window for Clarivate to calculate a normalized rate. The score is a decimal where 1.0 equals the world average for that category and document type: a JCI of 1.5 means the journal's papers are cited 50% more than the world average for comparable papers in that field, and a JCI of 0.5 means cited half as often.
Journal Registration (ISSN Registration)
Journal registration is the process of obtaining an International Standard Serial Number (ISSN) for a serial publication -- an 8-digit identifier assigned exclusively by the ISSN Network (a national ISSN Centre, or the ISSN International Centre for countries without one) that uniquely identifies a specific title in a specific medium. A publication counts as "registered" only once the relevant Centre has verified the title, publisher, and format and formally assigned it an ISSN through the ISSN Portal -- not merely because a publisher claims or advertises one. It is a bibliographic-identification step, distinct from peer-review vetting, DOAJ inclusion, database indexing, or any copyright/trademark filing.
KBART (Knowledge Bases And Related Tools)
A NISO Recommended Practice (currently RP-9-2014, KBART Phase II) that defines a common tab-delimited data format and field set for content providers to supply title-level holdings metadata — title, ISSN/ISBN, coverage dates, embargo, and URL — to knowledge-base vendors, so that link resolvers and discovery services can accurately determine and display full-text availability.
Kerndatensatz Forschung (KDSF)
The Kerndatensatz Forschung (KDSF) is Germany's national standard defining how universities and non-university research institutions structure and report information about research activity -- projects, personnel, publications, funding, and infrastructure -- using a shared base data model (Basisdatenmodell), topic-scoped modules, and standardised reference queries (Referenzabfragen). Voluntary adoption was recommended by the German Council of Science and Humanities (Wissenschaftsrat) in January 2016; ongoing governance now sits with the Kommission fur Forschungsinformationen in Deutschland (KFiD).
Kudos (research-impact platform)
Kudos is a free platform on which an author or institution takes a Crossref-DOI-registered publication and (1) attaches a plain-language explanation of the work, (2) enriches it with supporting links, data, or media, and (3) shares that package through built-in social/email distribution tools, with usage, download, and attention-metric data reported back afterward. It does not register DOIs, host full text, or index the literature itself -- it is an outreach and measurement layer sitting on top of publications already registered elsewhere.
LIMS vs LIS: What’s the Difference?
A system is a LIMS (Laboratory Information Management System) when its central unit of record is the sample: it registers specimens and aliquots, tracks them through testing, quality control and disposition or storage, manages instrument interfaces and inventory, and is built for research, industrial, environmental, quality-control and reference laboratories. A system is an LIS (Laboratory Information System, sometimes called a Laboratory Information Systems platform) when its central unit of record is the patient encounter: it receives test orders from an electronic health record (EHR), routes the associated specimen through a clinical diagnostic workflow, and returns a coded, clinician-facing result (commonly mapped to LOINC codes) back into the patient's chart, typically operating under CLIA and CAP oversight in the United States. The dividing line is audience and unit of record, not the feature list — both categories commonly include barcoding, instrument connectivity, chain-of-custody tracking and audit trails, so those features alone do not tell you which category a given system belongs to.
Link Resolver
A link resolver is OpenURL-based library-discovery software that accepts a context-sensitive citation link (from a database, discovery layer, or search engine), checks it against the library's knowledge base of current subscriptions and open-access holdings, and, if a match exists, routes the user to the specific full-text copy that institution has legitimate access to -- rather than to a fixed, publisher-assigned location. It is the software behind a library's "Find It", "Get It", or "Get Full Text" button.
Many-analysts study
A study design in which a single dataset and research question are given to multiple independent analysts or teams who proceed without coordination, and the distribution of their conclusions is then compared.
Metadata
Structured information that describes, identifies, or explains a research output — most often a dataset — independent of the content itself, so that the output can be found, correctly interpreted, cited, and reused without the data (or document, sample, or instrument) having to be opened or examined directly. In research data management, metadata is what a repository, catalog, or search index actually reads to determine whether a record matches a query; the underlying data file is opaque to that process until metadata points to it.
Metadata enrichment
A record's metadata is enriched when new descriptive, administrative, or relational elements are added to it after initial deposit or harvest — typically by resolving free-text values (author names, affiliation strings, subject terms) against authority files and controlled vocabularies (ORCID, ROR, funder registries, subject-classification schemes), or by pulling additional attributes from an external API (citation counts, open-access status, computed topic classifications). It is distinct from correcting or normalizing values already present (data cleaning) and from the original act of depositing the baseline record (initial deposit) — enrichment specifically adds elements the record did not previously have.
Metadata Format
A metadata format (also called a serialization format) is the concrete syntactic encoding used to write metadata values out as a machine-readable byte stream for storage, exchange, or transmission — for example XML, JSON, RDF/Turtle, or YAML. It answers 'how is this metadata physically written down,' which is a separate question from what a metadata schema answers: 'which fields exist and what do they mean.' A schema (Dublin Core, DataCite, Darwin Core, PROV-O) defines the field set and semantics; a format defines the byte-level syntax those fields get poured into. The same schema is routinely available in more than one format, and the same format is routinely used to carry many unrelated schemas — so 'metadata format' and 'metadata schema' are not interchangeable, even though the two terms get conflated constantly in casual usage.
Metadata Schema
A defined, published set of descriptive fields (elements or properties) — each with a name, definition, expected data type or controlled vocabulary, and obligation/cardinality rules — used to describe a resource consistently enough for people and systems to find, cite, exchange, and reuse it. Distinct from a metadata record, which is one specific instance of data filled into that specification.
Microsoft Academic Graph (MAG)
A large-scale, free scholarly knowledge graph built and maintained by Microsoft Research — publications, authors, affiliations, venues, and a machine-generated fields-of-study taxonomy, linked by citation and authorship edges, and distributed via the Microsoft Academic website, the Microsoft Academic Knowledge API, and periodic bulk data dumps — that Microsoft fully retired on December 31, 2021, and no longer maintains, updates, or serves in any form.
Multiverse analysis
An analytical approach in which all reasonable combinations of data-processing and modelling choices are executed, producing a distribution of results that displays the impact of researcher degrees of freedom on the conclusion.
Nature Portfolio Reporting Summary
A mandatory, standardized fillable-PDF form Nature Portfolio requires with manuscript submission for life sciences, behavioural & social sciences, or ecological/evolutionary/environmental sciences research articles, disclosing statistics reporting, software/code, data availability, and a field-specific study-design track (sample size, data exclusions, replication, randomization, blinding, or their track-specific equivalents), with negative disclosures required rather than fields left blank; the completed form is shared with editors/reviewers during assessment and published alongside the accepted paper.
Nextflow (concept)
A workflow orchestration system based on dataflow programming with a Groovy-based domain-specific language, designed for scalable, container-native, multi-platform execution of computational pipelines.
NIH Genomic Data Sharing (GDS) Policy
A study falls under NIH's Genomic Data Sharing (GDS) Policy (2014) when it is NIH-funded and generates large-scale human or non-human genomic data (GWAS, SNP arrays, WGS/WES, transcriptomic, epigenomic, metagenomic, or gene-expression data). Covered human-data studies must use GDS-specific prospective informed consent language, obtain an Institutional Certification from the awardee institution (via its Signing Official, in coordination with the IRB) confirming consistency with participant consent, and submit data to an NIH-designated controlled-access repository -- dbGaP for most human genomic data -- under a Data Use Certification Agreement. The GDS Policy predates and operates alongside, not inside, the broader 2023 NIH Data Management and Sharing (DMS) Policy.
NIH Rigor and Reproducibility policy
The set of US National Institutes of Health policies, effective from 2016, requiring applicants and grantees to address scientific premise, scientific rigour, biological variables (including sex as a biological variable), and authentication of key biological and chemical resources in grant applications.
NISO Access and License Indicators (ALI)
NISO RP-22: machine-readable metadata (free_to_read, license_ref) letting systems determine an article's access status and reuse terms without parsing prose.
NISO Institutional Identifier (I2)
A now-inactive NISO initiative (2008–2013) whose output, ANSI/NISO RP-17-2013, is a Recommended Practice for identifying organizations in the scholarly information supply chain. It did not define a new identifier syntax of its own; instead it specified how to apply the existing ISNI standard (ISO 27729) to institutions. The working group has not been active since publishing the Recommended Practice in 2013, and I2 has no ongoing registry of its own — in current research-infrastructure practice, ROR is the identifier institutions actually register for and that funders, publishers, and CRIS systems require.
NISO Recommended Practice (RP)
A consensus-based guidance document published by the National Information Standards Organization (NISO) that describes an accepted, emerging, or leading-edge method, format, or convention for information exchange. A Recommended Practice is developed through NISO's normal working-group and public-comment process but is approved without the formal ANSI ballot of NISO's Voting Members that a full ANSI/NISO Standard requires, and it is identified with the sequential designation RP-[number]-[year] (e.g., RP-9-2014) rather than the Z39.xx numbering used for Standards.
NSF FastLane (Legacy Grants Portal)
NSF FastLane (fastlane.nsf.gov) was the National Science Foundation's original web-based electronic system for proposal preparation, submission, and award administration, launched in the 1990s. A document or workflow is properly described as "a FastLane submission" or "a FastLane-era award" if it was prepared, submitted, or administered through that platform prior to NSF's cutover to Research.gov. As of 2023, FastLane is decommissioned as a live proposal-submission channel: NSF's Proposal & Award Policies & Procedures Guide (PAPPG) NSF 23-1 removed FastLane as a submission option for all NSF funding opportunities effective January 30, 2023, and NSF phased out FastLane's remaining functions (award documents, organizational reports, Grants.gov integration) in the months that followed, including a further phase in November 2023. The fastlane.nsf.gov domain now 301-redirects to research.gov, confirming there is no live FastLane instance to submit to. Research administrators encountering the term today are almost always dealing with historical records, legacy award documentation, or institutional process guides that have not yet been updated to reflect the transition.
OAIS Reference Model (ISO 14721)
A conceptual framework, standardised as ISO 14721, that defines the functions, information packages, and terminology an archive needs to preserve digital (or physical) information and keep it accessible to a defined community over the long term, independent of the specific technology used to implement it.
ONIX (ONline Information eXchange)
An XML-based family of metadata standards, maintained internationally by EDItEUR (with regional groups Book Industry Communication/BIC in the UK and BISG in the US), used to exchange structured product and rights information across the publishing supply chain. A metadata record qualifies as ONIX when it validates against one of EDItEUR's published ONIX XML schemas/DTDs (the current major version for book-trade metadata is ONIX for Books 3.0, extended by 3.1) and carries the standard's defined composite structure, code lists, and identifiers rather than a proprietary or free-text feed.
Open archive
A repository that is compliant with the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) and therefore facilitates the sharing of metadata for a variety of purposes, most notably the compilation tasks performed by aggregator databases.
Open Database License (ODbL)
The Open Database License (ODbL) is a copyleft license from Open Data Commons (a project of the Open Knowledge Foundation), version 1.0 published June 2009, that grants reuse rights to a database subject to attribution and share-alike conditions, drafted specifically around database rights and database structure/contents rather than adapted from a creative-works copyright license.
Open Science
The umbrella movement and set of practices for making every stage of the research process — publications, data, code, methods, and materials — openly accessible, transparent, and reusable, rather than treating openness as something that applies only to the final published article. UNESCO’s 2021 Recommendation on Open Science, the first international normative instrument on the subject, organizes it into four pillars: open scientific knowledge, open science infrastructures, open engagement of societal actors, and open dialogue with other knowledge systems.
Open Science Badges (Open Data, Open Materials, Preregistered)
The Open Science Badges are a set of journal-awarded visual icons, created and administered by the Center for Open Science (COS), that certify a published article met a disclosed openness standard for one of three practices: publicly posting the study's underlying data (Open Data), publicly posting the materials needed to reproduce the reported procedure and analysis (Open Materials), or registering the study design -- and, for the Preregistered+Analysis Plan variant, the analysis plan -- in a public registry before data collection began (Preregistered). A journal awards a badge either on a simple author self-disclosure statement or after independent editorial verification, and displays the badge icon alongside the article together with a link to the disclosed data, materials, or registration.
Open Science Framework (OSF)
A free, open-source web platform operated by the nonprofit Center for Open Science (COS) that lets a research team manage a project's full lifecycle from one persistent, citable project page: file storage and version control, collaborator permissions, preregistration of hypotheses and analysis plans, and public archiving with a DOI. 'Open Science Framework' is often shortened or misremembered as 'Open Science Foundation' — the organization that builds and runs OSF is the Center for Open Science, not an 'Open Science Foundation'; no organization by that name operates this infrastructure.
OpenAIRE EXPLORE
The end-user-facing discovery service of OpenAIRE that allows researchers, funders, and the public to search and browse the OpenAIRE Graph by publications, datasets, software, projects, organisations, and funders.
OpenAIRE Graph
An open scholarly knowledge graph maintained by OpenAIRE that aggregates and links publications, datasets, software, projects, organisations, and people harvested from thousands of repositories, journals, and CRIS systems across Europe and globally, with deduplication, enrichment, and link inference applied.
OpenAIRE Nexus
An OpenAIRE-led Horizon-Europe-funded initiative (2021-2024) that bundles OpenAIRE services and tools into a coherent service portfolio for delivery through the European Open Science Cloud (EOSC), targeting researchers, content providers, research communities, funders, and policy makers.
Operational Definition
A statement that specifies the exact, observable procedure used to measure or manipulate an abstract construct or variable within a specific study, so that another investigator could apply the same procedure to the same phenomenon and obtain comparable data. An operational definition is distinct from a conceptual (or constitutive) definition, which states what a construct means in theory without specifying how it will be observed; a definition only counts as operational if it names the instrument, coding rule, threshold, or manipulation used, not just the general idea being studied.
ORCID Featured Works
A record is using ORCID Featured Works when the iD holder has designated one or more of their public works (up to a maximum of five) to appear in a dedicated "Featured" section at the top of the Works list on their ORCID record, each marked with a star icon. Featuring is a curation layer applied on top of an already-public work: only works whose visibility is already set to "Everyone" are eligible to be featured, and a work is automatically removed from the Featured list the moment its visibility is changed away from public. As of verification, the feature applies to the Works section only — there is no equivalent "Featured" mechanic for Employment, Education, Funding, Peer Review, or Distinctions entries on the record.
ORCID iD for NIH Senior/Key Personnel
The ORCID iD requirement for NIH Senior/Key Personnel means: every individual designated as Senior/Key Personnel on an NIH grant application — the Program Director/Principal Investigator (PD/PI) plus any other individual who contributes to the scientific development or execution of a project in a substantive, measurable way, regardless of whether they receive salary support from the award — must (1) hold a registered ORCID iD, (2) link it to their own eRA Commons Personal Profile account, an action NIH requires the individual to complete themselves rather than delegate to grants administration staff, and (3) have that iD populate correctly in the Persistent Identifier (PID) field of the NIH Common Form versions of the Biographical Sketch and Current and Pending (Other) Support documents, both generated through SciENcv. This applies to applications with due dates on or after January 25, 2026 (NIH Guide Notice NOT-OD-26-018). An application missing a required, correctly linked ORCID iD for any listed Senior/Key Personnel is not merely flagged — as of May 8, 2026, eRA Commons system validations treat this as a hard error that blocks submission outright (NOT-OD-26-079), after an initial warning-only enforcement period (extended once, via NOT-OD-26-033, through May 7, 2026).
ORCID invited position
An affiliation item in an ORCID record asserting that the iD holder holds or held a paid or unpaid invited role — such as a visiting lectureship, honorary or guest-researcher appointment, or emeritus professorship — at a named organisation, as distinct from an ordinary employment relationship. Captured with the organisation's name and disambiguated identifier (typically a ROR ID), department, role/title, start date, and optional end date. In the ORCID record schema and API, "invited-position" is one of seven distinct affiliation-type values (alongside employment, education, qualification, distinction, membership, and service) — it is a role the person actually holds, not merely an honour received (that is "distinction") and not professional service such as board or committee work (that is "service"), even though ORCID's own web entry form groups invited position and distinction into one combined section.
ORCID service
An affiliation entry in an ORCID record documenting an ongoing donation of a researcher's time or expertise to an organization or community — board, committee, or editorial-board membership, review panels, or extension work — captured with a role title, disambiguated organization identifier, and start/end dates.
ORCID Sign-In / Authentication
ORCID sign-in ("Sign in with ORCID") is the OAuth 2.0/OpenID Connect process by which a third-party system (a journal, funder, or institutional platform) redirects a researcher to orcid.org to authenticate, then receives back an authenticated ORCID iD and, if broader scopes were requested and approved, an access token to read or write specific parts of the researcher's ORCID record. It is distinct from simply holding an ORCID iD: an iD is a registered identifier; signing in is the act of a specific external system verifying, at that moment, that the researcher controls it, and optionally being granted ongoing record access.
Organization unit
In CRIS terminology, an entity representing a structural component of a research organisation — a faculty, school, department, institute, centre, lab, or research group — with its own identifier, name, parent and child relationships, type, start and end dates, and links to People, Projects, and Outputs.
OSF Preprints
<p>A work counts as an <strong>OSF Preprints</strong> deposit if it is hosted on the preprint infrastructure built and operated by the <a href='https://www.cos.io' rel='nofollow'>Center for Open Science (COS)</a> -- either on the generalist osf.io/preprints server itself, or on one of the branded, discipline-specific community servers (PsyArXiv, SocArXiv, EarthArXiv, and others) that run on the same underlying OSF backend but carry their own name, editorial moderators, and disciplinary scope. It is distinct from an OSF <em>project</em> (the private/public research-management workspace) and from OSF Registries (preregistrations): a preprint is a discrete, versioned, publicly posted manuscript that receives its own Crossref-registered DOI, separate from any OSF project it may have originated in.</p>
P-hacking
The practice of selectively reporting or adjusting analytical choices in order to obtain a statistically significant p-value, typically below the conventional 0.05 threshold.
PDF/A (ISO 19005)
A file is PDF/A-conformant only when it validates against one of the four parts of ISO 19005: every font it uses is embedded rather than referenced, color is specified in a device-independent way, and the file contains no encryption, no JavaScript or executable content, no audio or video, and no reference to external content the renderer would need to fetch. These restrictions make the file self-contained and reproducible without depending on the specific software environment that created it.
Pending publication (Crossref)
A Crossref record type that lets a publisher (a Crossref member) register a DOI and a minimal metadata record -- at minimum the member name, journal title, and the manuscript's accepted date -- for a manuscript that has already cleared peer review and received a formal editorial acceptance decision, but has not yet been published online. Until the work is formally published, the DOI resolves to a Crossref-hosted landing page carrying an acceptance notice and the deposited metadata, not to the article itself. The publisher must later update that same DOI's metadata record to a full journal-article registration once the work is published; only then does the DOI resolve directly to the published content. A record qualifies as a pending publication only if it was deposited by the eventual publisher for an already-accepted manuscript -- an unrefereed preprint, or a manuscript still under review, does not qualify.
Persistent data identifier
A persistent identifier (typically a DataCite DOI, but also IGSN for samples, Handle, ARK, or other PID-scheme identifier) assigned to a research dataset to support stable citation, attribution, and resolution over the long term.
Persistent Identifier (PID)
A unique, actionable, long-term identifier assigned to a single entity in the research ecosystem -- a work, a person, an organisation, a piece of equipment, a project, or a dataset -- that is registered with a scheme-specific provider, stays associated with that entity even as its location or attributes change, resolves (typically over HTTP) to current metadata or the resource itself, and is designed to remain valid for decades, well beyond the lifetime of any one repository, publisher, or institutional system.
Person record (CRIS)
In CRIS terminology, the entity representing an individual person involved in research — researcher, RA, PhD candidate, or other contributor — with metadata including names (preferred and historical), identifiers (ORCID iD, ISNI, local HR ID), employments, qualifications, and links to outputs, projects, and organisational units.
PICOTS Criteria
A research question is PICOTS-structured when it explicitly specifies all six elements — Population, Intervention, Comparator, Outcome, Timing, and Setting — rather than only the four PICO elements, typically to derive study-eligibility criteria for a systematic review, evidence synthesis, or health-technology assessment protocol.
PREMIS
A resource's metadata qualifies as PREMIS metadata when it is structured according to the PREMIS Data Dictionary for Preservation Metadata (currently version 3.0, maintained by the Library of Congress), populating semantic units under one or more of PREMIS's four in-scope entities — Object, Event, Agent, and Rights — to document what a digital object is, what has happened to it over time, who or what acted on it, and what permissions govern its preservation. PREMIS is a data dictionary, not a file format or package standard in the way METS is; it is most commonly implemented as an XML schema (premis.xsd) embedded inside a METS document's administrative metadata section, or serialized natively inside a repository's own preservation system.
PRISMA 2020
The 2020 update of the Preferred Reporting Items for Systematic reviews and Meta-Analyses, a 27-item checklist with accompanying flow diagram for reporting systematic reviews.
PRISMA Flow Diagram
The standardized figure PRISMA 2020 uses to report the flow of records through a systematic review or meta-analysis, organized under three labeled phases (Identification, Screening, Included), with a numeric box for every stage of attrition and a stated reason for every full-text exclusion.
PRISMA-ScR (PRISMA Extension for Scoping Reviews)
A report is PRISMA-ScR compliant when it satisfies PRISMA-ScR's 20 essential reporting items (Tricco et al. 2018), frames its question in PCC (Population, Concept, Context) rather than PICO terms, and documents evidence "charting" and mapping rather than risk-of-bias-weighted synthesis -- the checklist purpose-built for scoping reviews, distinct from the core 27-item PRISMA 2020 checklist used for systematic reviews of intervention effects.
Project (CRIS)
In CRIS terminology, an entity representing a discrete research project with a defined scope, time period, participating people and organisations, funding source(s), and intended or actual outputs — typically the central organising entity for activity-level reporting.
PROSPERO Registration
PROSPERO registration is the prospective submission of a systematic review's protocol (research question, PICO, search strategy, eligibility criteria, and analysis plan) to the International Prospective Register of Systematic Reviews (PROSPERO), maintained by the Centre for Reviews and Dissemination (CRD) at the University of York and funded by the UK's NIHR, before the review has progressed beyond data extraction. Eligible reviews must have a health-related outcome; a registered record receives a permanent, publicly citable CRD42-prefixed identifier.
protocols.io (concept)
protocols.io is a free, open platform (owned by Springer Nature since 2023) for developing, versioning, and sharing detailed step-by-step research method protocols across life-science, physical-science, and computational disciplines. Each published protocol receives a DataCite DOI, so a specific, citable version of a method persists even as the protocol is later revised, forked, or optimized by other labs.
PROV-O (Provenance Ontology)
PROV-O is the W3C's OWL2 ontology (a W3C Recommendation since 30 April 2013) for expressing data provenance as machine-readable RDF: it defines Entity (a thing with fixed aspects), Activity (something that acts upon or generates entities over time), and Agent (something responsible for an activity or entity), connected by properties such as wasGeneratedBy, used, wasAssociatedWith, and wasDerivedFrom.
Publish or Perish (Harzing’s citation-analysis software)
Publish or Perish is a free desktop program, created by Anne-Wil Harzing and first released in October 2006, that queries academic data sources (as of version 8: Google Scholar, Google Scholar Profile, Crossref, PubMed, OpenAlex, Lens.org, Scopus, Semantic Scholar, and Web of Science) for a given author, journal, or search term, then calculates citation-based metrics such as the h-index and g-index from whatever results are returned. It holds no citation data itself and is not a database, persistent identifier registry, or CRIS -- it is a query-and-calculate client for other providers' data, so results vary by which source was queried and change over time as underlying citation counts change.
Published Data
<p>Data is 'published' when it has been formally released as a discrete, citable research output: deposited in a repository or registry that assigns it a persistent identifier (typically a DOI), described with metadata sufficient for independent discovery and reuse (at minimum the DataCite mandatory set -- Identifier, Creator, Title, Publisher, PublicationYear, ResourceType), and made accessible under a stated licence. A dataset that is merely stored, backed up, or informally emailed to a collaborator is not 'published' in this sense, even if the underlying files are identical -- publication is a formal act with a fixed citation form, not a description of where the bytes happen to live.</p>
Pure
A commercial CRIS product, originally developed by Atira A/S in Denmark and now owned by Elsevier, used by universities and research organisations to manage publications, projects, people, organisational units, awards, equipment, and external engagement, and to expose this information through a configurable public 'research portal'.
QDR (Qualitative Data Repository)
The Qualitative Data Repository (QDR) is a domain repository dedicated to curating, preserving, and providing access to digital data generated through qualitative and multi-method social science research — interview transcripts, fieldnotes, focus-group recordings, participant-observation records, and other unstructured or semi-structured data types that do not fit the tabular, quantitative deposit model most general-purpose repositories are built around. It is hosted by the Center for Qualitative and Multi-Method Inquiry at Syracuse University's Maxwell School of Citizenship and Public Affairs. A dataset qualifies as a QDR deposit when it is submitted through QDR's own curation and review workflow, receives a persistent identifier (a DataCite DOI), and is described with QDR's qualitative-data-specific metadata and, where the depositing researcher chooses to use it, QDR's Annotation for Transparent Inquiry (ATI) format for linking published claims back to underlying data excerpts.
Randomized Controlled Trial (RCT)
A prospective clinical study design in which participants are allocated to two or more comparison groups purely by chance (randomization), with at least one group — the control arm — receiving a placebo, an active comparator, or standard-of-care rather than the intervention under investigation. Randomization is what distinguishes an RCT from an observational study; the presence of a control arm is what distinguishes it from a single-arm trial. Blinding (masking) of participants, investigators, or outcome assessors is a common but analytically separate design feature — an RCT can be randomized, controlled, and still unblinded (open-label).
Re3data (concept)
Registry of Research Data Repositories: a global registry, operated by DataCite and partner institutions, that lists research data repositories worldwide with descriptive metadata about their disciplines, content types, access conditions, and policies, helping researchers locate suitable repositories for deposit and discovery.
REDCap (Research Electronic Data Capture)
REDCap (Research Electronic Data Capture) is a secure, metadata-driven, web-based software platform for building surveys and case-report-form databases for research, originated at Vanderbilt University in 2004. An instance only counts as REDCap if it runs the actual Vanderbilt-distributed software, licensed at no cost to non-profit institutions through the REDCap Consortium and installed/administered by that partner institution -- there is no direct commercial purchase path or self-service individual signup. Typical uses include participant surveys, longitudinal case report forms, and multi-site clinical data capture, frequently chosen specifically because the platform supports HIPAA, FDA 21 CFR Part 11, FISMA, and GDPR compliance controls at the institutional level.
Reporting Guidelines
<p><strong>Reporting guidelines</strong> are checklists, flow diagrams, or structured-text tools that specify the minimum set of information authors must include when writing up a specific type of study, so that readers, peer reviewers, and editors can judge what was done and trust that nothing material has been left out. The EQUATOR (Enhancing the QUAlity and Transparency Of health Research) Network, an international initiative founded in 2006 and based at the University of Oxford's Centre for Statistics in Medicine (NDORMS), is the field's central clearinghouse: it defines a reporting guideline as a checklist, flow diagram, or structured text developed using an explicit, documented consensus methodology to guide authors in reporting a specific type of research, and it maintains a searchable library of several hundred such guidelines spanning study designs, specialties, and report sections.</p><p>A reporting guideline is not a methodological standard. It does not tell an investigator how to design a trial, calculate a sample size, randomize participants, or run an assay -- that is the job of methodological/conduct standards such as ICH Good Clinical Practice, CIOMS guidance, or a funder's data-collection protocol. Nor is a reporting guideline a critical-appraisal or risk-of-bias tool -- it does not score how well a study was conducted (that is the role of instruments such as the Cochrane risk-of-bias tools). A reporting guideline answers a narrower, prior question: <em>given that the study was conducted a particular way, has the write-up disclosed enough about it -- the eligibility criteria, the randomization method, the funding source, the flow of participants -- for a reader to evaluate and, ideally, reproduce the work?</em> An author can follow a reporting guideline to the letter while having run a methodologically weak study, and a well-conducted study can still be reported so incompletely that a reader cannot tell.</p>
Repository
Repositories preserve, manage, and provide access to many types of digital materials in a variety of formats.
Repository Citation
A repository citation is the reference used to attribute and locate a dataset (or other research output such as software or a physical sample) deposited in a data repository, built around a persistent identifier — typically a DOI minted via DataCite — rather than around the journal, volume, and page numbers used for an article citation. A citation counts as a repository citation when it points directly to the deposited data object itself (resolving through its own persistent identifier to the repository landing page or the data), not to a journal article, data paper, or other publication that merely describes or analyzes that data.
Research activity (CRIS)
In CRIS terminology, an entity representing a research undertaking — typically a project, programme, or organised research effort — with its own start and end dates, participants, funding sources, outputs, and host organisation, around which CRIS data accretes over the activity's lifetime.
Research Data
Research data is the recorded, factual material generated or collected in the course of a research project that is needed to validate, reproduce, or build on that project's findings &mdash; raw measurements, observations, instrument readings, survey responses, code outputs, and other structured or unstructured records, regardless of medium or discipline. In U.S. federally funded research specifically, the term carries a narrower, load-bearing regulatory meaning under 2 CFR &sect; 200.315(e)(3) (the Uniform Guidance, which consolidated and replaced the older OMB Circular A-110 in 2014): research data means <em>'the recorded factual material commonly accepted in the scientific community as necessary to validate research findings.'</em> That specific definition is what determines what a federal award recipient must make available in response to a Freedom of Information Act (FOIA) request, and it is the baseline many institutions use to scope what a <a href='/dictionary/term/data-management-plan-dmp'>data management plan (DMP)</a> is actually obligated to describe. Distinguishing 'research data' in this operational sense from the much larger set of everything a researcher produces during a project &mdash; drafts, correspondence, physical samples, unfinished analysis &mdash; is the first practical step in scoping a DMP, a repository deposit, or a records-retention schedule, and underpins <a href='/pillar/rdm'>research data management (RDM)</a> practice generally.
Research Data Alliance (RDA)
The Research Data Alliance (RDA) is a global, community-driven organisation, launched in March 2013 by the European Commission, the US National Science Foundation, the US National Institute of Standards and Technology, and the Australian government, to build the social and technical infrastructure needed for open data sharing across disciplines, technologies, and national borders. RDA operates chiefly through two kinds of groups: time-bounded Working Groups, which typically run around 18 months and produce concrete, adoptable deliverables, and open-ended Interest Groups, which coordinate ongoing discussion, surveys, and less formal outputs across a topic area. Formally endorsed Working Group deliverables are published as RDA Recommendations; less formal outputs from any group type are published as Supporting Outputs. Something qualifies as "RDA" only when it refers to the organisation itself — its membership, governance (Council, Technical Advisory Board, Secretariat), plenaries, and working/interest-group structure — not to any single recommendation, schema, or standard an RDA group has produced, which are distinct, separately citable outputs.
Research entity (CRIS)
A first-class object in a CRIS data model — typically Person, Project, Publication, Organisation Unit, Funding, Equipment, or Activity — that has its own identifier, metadata schema, and relationships to other entities, and that can be managed and reported on independently.
Research output (CRIS)
In CRIS terminology, an entity representing a discrete product of research — a publication, dataset, software release, patent, performance, exhibition, or other recognised output — recorded with its own identifier, type, date, contributors, and relationships to people, organisations, and projects.
Research Transparency
Research transparency is the practice of making the specific decisions, materials, data, and code behind a study visible and checkable at the point they are used, not the broader movement to open up research outputs generally (that is <a href='/dictionary/term/open-science'>open science</a>). A study demonstrates research transparency to the degree it discloses, and ideally shares and cites, what it did: whether hypotheses and analysis plans were specified before or after seeing the data, where materials and data can be obtained, whether analytic code is available, and whether the design and reporting follow a recognized standard. The most widely adopted operational framework for this is the Transparency and Openness Promotion (TOP) Guidelines, published by the Center for Open Science (COS), which turns 'be transparent' into eight checkable standards — citation, data transparency, analytic methods (code) transparency, research materials transparency, design and analysis transparency, study preregistration, analysis plan preregistration, and replication — each ratable at one of three levels of stringency, from disclosure through requirement to independent verification.
Researcher degrees of freedom
The decisions an analyst makes during a study (inclusion criteria, outcome definition, model specification, covariate set) any of which, if made differently, would yield a different result.
Researcher Identifier
A persistent, unique digital identifier assigned to an individual person -- as distinct from an identifier for a work, an institution, or a grant -- that disambiguates that researcher across name variants, career moves, and multiple scholarly systems, and that is designed to stay resolvable and machine-readable for the researcher's whole career.
Researcher Profile Optimization Study (RPOS)
An in-progress Federal Demonstration Partnership (FDP) initiative, run under FDP's Open Government: Research Administration Data (OG:RAD) subcommittee, that surveys researchers and research-administration offices across FDP's member institutions to document how federally required researcher profile information -- biosketches, current & pending support, and collaborator disclosures -- is captured across agencies and platforms, and to recommend changes that reduce duplicate data entry and support interoperability, research security, and compliance.
Researcher webpage
A webpage featuring a researcher's profile, which possibly may also provide links to their publications.
ResearcherID
ResearcherID is Clarivate's proprietary author identifier for the Web of Science ecosystem: a unique alphanumeric code (format letter(s)-numbers-year) assigned when a researcher creates a Web of Science Researcher Profile, used to disambiguate that person's publications and citation metrics within Web of Science and InCites. Since Clarivate retired Publons as a separate product in April 2022, ResearcherID exists as the identifier component of the unified Web of Science Researcher Profile rather than a standalone product, and (once a researcher explicitly connects the two accounts) two-way syncs publications with a connected ORCID iD.
Rigor and Transparency Index (RTI)
The Rigor and Transparency Index (RTI) is the average SciScore — an automated score of methods-rigor and resource-identification reporting — across a journal's or institution's articles in a given year, published by SciCrunch as a benchmarking metric for reporting quality. It is calculated independently of citation-based metrics such as Journal Impact Factor and is meant to be compared only within the same field or subfield.
RIM
Research Information Management: the organisational practice and the supporting systems and processes by which a university or research organisation collects, manages, and uses information about its research activities, encompassing both the technical CRIS layer and the people and policies around it.
Robustness
The stability of a study's conclusions under reasonable variations in analytical choices, model specification, sample inclusion, or measurement, on the same data.
Robustness check
An additional analysis, supplementary to the headline result, that varies one or more analytical choices in order to demonstrate that the main conclusion is not artefactual to those choices.
RRID (Research Resource Identifier)
A string of the form RRID: followed by an accession number issued by an authoritative source database (the Antibody Registry for AB_ numbers, Cellosaurus for CVCL_ cell-line codes, MGI/ZFIN/RGD/IMSR for model organisms, or SciCrunch's own SCR_ registry for software and databases) that resolves, via the SciCrunch resolver or a compatible resolver such as n2t.net, to a single curated record describing exactly one research resource. Developed through FORCE11's Resource Identification Initiative and maintained via SciCrunch's Resource Identification Portal, an RRID exists specifically so a research resource used in a study's methods can be cited unambiguously and traced across the literature, the way a DOI identifies a publication or an ORCID iD identifies a researcher.
RTSM (Randomization and Trial Supply Management)
A system counts as RTSM (Randomization and Trial Supply Management) -- also called IRT (Interactive Response Technology) -- if, within a single platform used by trial sites, it performs at least one of: (1) subject randomization to a treatment arm per the protocol's randomization scheme at enrollment, with blind maintenance and an auditable emergency-unblinding path where the trial is blinded; or (2) trial-supply management -- real-time kit- or dose-level inventory tracking and automatic resupply triggering to sites or depots. Most production RTSM/IRT platforms do both together. A platform that only captures clinical data (an EDC) or only manages site financial, regulatory, and monitoring workflow (a CTMS) is not RTSM/IRT, even though all three are routinely used side by side, and sometimes integrated, on the same trial.
Sample repository
A repository for physical research samples — geological, environmental, biological, or material — that catalogues, stores, and provides access to samples for downstream analysis, often issuing persistent identifiers (IGSN, DataCite DOI) for citation and provenance tracking.
Sampling Methods: Probability and Non-Probability Types Explained
A sampling method is the defined procedure a study uses to select which members of a population are observed or measured. It qualifies as a probability method only if every population member has a known, non-zero chance of selection via a defined chance mechanism (simple random, stratified, systematic, or cluster sampling) -- this is what supports formal statistical inference and generalizability claims back to the population. Non-probability methods (convenience, purposive/judgmental, snowball, quota sampling) select participants without a defined chance mechanism, which is legitimate and standard in qualitative and exploratory research but does not, on its own, support formal population-level generalization.
Scientific Data Management System (SDMS)
A scientific data management system (SDMS) is laboratory software that automatically captures, indexes, and archives electronic data files generated by lab instruments (chromatography, spectroscopy, balances, plate readers, raw instrument logs) in their original native format, and makes that archive centrally searchable, version-controlled, and audit-trail-protected. It is distinguished operationally from a shared drive or manual file store by automated unattended capture direct from the instrument, metadata extraction that makes files findable without opening each one, and a tamper-evident audit trail logging every access, export, or modification attempt. An SDMS is the raw-data archive layer beneath a LIMS (sample/workflow/results management) and an ELN (experimental narrative record), and typically integrates with one or both rather than replacing either.
Scientific rigour
The strict application of the scientific method to ensure unbiased and well-controlled experimental design, methodology, analysis, interpretation, and reporting of results.
SciScore
SciScore is an automated text-mining tool from SciCrunch that scores a manuscript's methods section (0-10) against named rigor and resource-transparency criteria -- randomization, blinding, sample-size justification, sex as a biological variable, and RRID-identified research resources -- drawn from NIH rigor guidelines, MDAR, ARRIVE, CONSORT, and STAR Methods, integrated into journal submission workflows at publishers such as the American Heart Association, Rockefeller University Press, and FASEB.
SciVal
A subscription research-analytics platform, owned and sold by Elsevier, built on Scopus abstract and citation data, that lets research institutions, funders and government bodies benchmark research performance, track collaboration, and analyse societal and funding impact at the level of a researcher, department, institution, country or custom peer group. It is a reporting and benchmarking layer on top of Scopus rather than a separate bibliographic database of its own -- everything SciVal shows is ultimately traceable back to a Scopus record.
Scopus
Scopus is Elsevier's proprietary, subscription-based abstract-and-citation database of peer-reviewed journal articles, conference proceedings, books, and book chapters, launched in November 2004. A title is "in Scopus" only once Elsevier's Content Selection and Advisory Board has evaluated and accepted it against criteria including peer-review policy, publication ethics, and editorial diversity — inclusion is an ongoing editorial decision, not automatic, and previously indexed titles can later be suspended or discontinued if they stop meeting those criteria.
Scopus Author ID: How to Find & Fix Yours
A Scopus Author ID is a numeric identifier that Elsevier's Scopus database automatically generates and attaches to a researcher's profile once that researcher has two or more documents indexed in Scopus — it is algorithmically assigned from article metadata (name variants, affiliation, subject area, co-authors, citation patterns), not requested or self-registered by the researcher, and a given person can end up with more than one Scopus Author ID if the algorithm fails to match all of their work to a single profile.
SCORE (Systematizing Confidence in Open Research and Evidence)
A DARPA-funded research program, run by DARPA's Defense Sciences Office from 2019 to roughly 2022, that tested whether quantitative 'confidence scores' -- generated from human forecasting/prediction markets, expert surveys, and automated algorithms -- could reliably estimate how likely a published social and behavioral science (SBS) claim was to replicate. The program sampled around 3,000 research claims from roughly 60 journals published between 2009 and 2018, had forecasters and algorithms score them before any new data collection, then commissioned independent replications on a subsample to check which forecasting method came closest to the actual outcome. It is now complete; DARPA lists it as inactive and retained for reference.
Secure Data Enclave
A controlled-access computing environment in which approved researchers analyze restricted or sensitive data in place, without the ability to download, copy, or otherwise export the raw underlying records. Access is typically granted only after training, credentialing, and a signed data use agreement, computation happens on infrastructure the data steward controls (an on-site terminal room, a remote-access session, or an isolated cloud workspace), and only aggregate or disclosure-reviewed outputs — tables, statistics, model results — are allowed to leave the environment. The defining feature is not where the data sits but that raw data never crosses the boundary; only vetted outputs do.
shortDOI
A shortDOI is a redirect alias, in the form 10/abcde, generated by the International DOI Foundation's shortDOI Service (shortdoi.org) for an existing, already-registered DOI name. It is not a separately registered persistent identifier: it resolves through the same DOI Handle System and https://doi.org/ resolver as the underlying DOI, is issued only once per DOI (requesting a shortDOI for a DOI that already has one returns the existing alias rather than minting a new one), and carries none of a DOI's registration guarantees on its own -- the original DOI remains the identifier of record for citation, archiving, and metadata purposes.
Snakemake (concept)
A Python-based workflow management system that expresses computational pipelines as rules with explicit inputs, outputs, and shell or script bodies, and infers a directed acyclic graph (DAG) of jobs from those rules.
SocArXiv
<p><strong>SocArXiv</strong> is a free, open-access preprint repository for the social sciences. It was founded in 2016 by sociologist Philip N. Cohen in partnership with the nonprofit <a href='/guides/center-for-open-science'>Center for Open Science (COS)</a> and runs on the <a href='/dictionary/term/open-science-framework-osf'>Open Science Framework (OSF)</a> &mdash; the same underlying infrastructure that hosts <a href='/dictionary/term/osf-preprints'>OSF Preprints</a> and sibling branded servers such as <a href='/dictionary/term/psyarxiv'>PsyArXiv</a> (psychology) and engrXiv (engineering). SocArXiv launched alongside those two services on December 5, 2016, as part of COS's branded-server model, in which a scholarly community operates its own named, separately moderated preprint service on shared OSF backend infrastructure rather than building hosting from scratch.</p> <p>SocArXiv is not a COS-run editorial operation in the way a journal is. It is governed by a volunteer steering committee of scholars and library professionals under the banner "SocOpen," and since 2021 its institutional home has been the University of Maryland Libraries, which provides the administrative and legal footing for the service as an ongoing academic unit rather than a standalone nonprofit.</p> <h2>Operational definition</h2> <p>A deposit is properly identified as a "SocArXiv preprint" if it meets all of the following:</p> <ul> <li>It was submitted through the SocArXiv service on OSF (osf.io/preprints/socarxiv), not simply cross-listed or mirrored from another repository such as <a href='/guides/arxiv-preprints-what-they-are-and-how-to-use-them'>arXiv</a> or SSRN.</li> <li>It falls within sociology or a closely related social-science field &mdash; SocArXiv accepts working papers, preprints, published-paper copies, data, and code across the social sciences broadly, not sociology alone.</li> <li>It has passed SocOpen's moderation screen, which is a scope-and-completeness check, not peer review: moderators confirm the item is scholarly, falls within a supported research area, is plausibly categorized, is correctly attributed to its stated authors, is in a moderated language, and is submitted in a text-searchable format (PDF or DOCX, not a scanned image).</li> <li>It has received an automatically assigned, persistent <a href='/dictionary/term/doi'>DOI</a> on upload, making it citable independent of eventual journal publication; a prior DOI issued by another publisher can also be linked to the same record.</li> <li>It carries one of SocArXiv's permitted reuse terms &mdash; CC BY 4.0, CC0 (public-domain dedication), or no license &mdash; selected by the author at deposit.</li> </ul> <h2>What SocArXiv is not</h2> <p>Posting to SocArXiv is explicitly not peer review, and a SocArXiv record is not, by itself, evidence that a manuscript has been formally accepted anywhere. Moderation checks scope and completeness, not scientific merit, methodology, or novelty. SocArXiv is also free to authors and readers &mdash; there is no submission fee or paywall, which distinguishes it from many venues that require an <a href='/dictionary/term/article-processing-charge'>article processing charge (APC)</a> at the journal stage.</p> <h2>Examples</h2> <ul> <li>A sociologist deposits a working paper analyzing survey data on labor-market outcomes to SocArXiv ahead of submitting it to a peer-reviewed journal, receiving a citable DOI immediately and using that DOI in grant reports and CVs while the manuscript is still under review elsewhere.</li> <li>A research team posts the accepted-but-not-yet-typeset version of a political-science article to SocArXiv under a publisher-permitted embargo, satisfying a funder's open-access mandate without waiting for the journal's own publication date.</li> </ul> <h2>Counter-example</h2> <p>A working paper in economics or sociology posted only to SSRN, a university's own institutional repository, or a personal website is <em>not</em> a "SocArXiv preprint," even though it may cover an identical topic and serve a similar function &mdash; the term refers specifically to material submitted through and moderated by the SocArXiv service on OSF, not to social-science preprints in general (for that broader category, see <a href='/dictionary/term/preprint'>preprint</a>).</p> <h2>Related terms</h2> <ul> <li><a href='/dictionary/term/preprint'>Preprint</a></li> <li><a href='/dictionary/term/osf-preprints'>OSF Preprints</a></li> <li><a href='/dictionary/term/psyarxiv'>PsyArXiv</a></li> <li><a href='/dictionary/term/open-science-framework-osf'>Open Science Framework (OSF)</a></li> <li><a href='/dictionary/term/doi'>DOI</a></li> <li><a href='/guides/how-to-choose-a-preprint-server'>How to Choose a Preprint Server</a></li> <li><a href='/guides/preprint-servers-explained'>Preprint Servers Explained</a></li> </ul>
Social Sciences Citation Index (SSCI)
A journal is considered SSCI-indexed when Clarivate's independent, in-house editorial reviewers have evaluated and accepted it into the Social Sciences Citation Index — the social-sciences-scoped citation index within the Web of Science Core Collection — based on Clarivate's quality criteria (peer review, editorial rigor, timeliness, ethical publishing standards) and its own separate impact/citation-influence assessment. SSCI inclusion is evaluated independently of the Science Citation Index Expanded (SCIE) and the Arts & Humanities Citation Index (AHCI): a journal is not automatically SSCI-indexed just because it appears in one of those companion indexes, and indexing is not permanent — a journal can later be suspended or delisted if it stops meeting Clarivate's criteria.
Software Heritage archive
A non-profit international initiative based at Inria that systematically crawls, archives, and preserves the world's publicly available source code, including its full version-control history, and issues persistent identifiers (Software Hash Identifiers, SWHIDs) to every archived artefact.
Specification curve
An analytic and visual technique that plots the estimated effect across a large set of theoretically defensible model specifications, ordered by effect size, to convey the sensitivity of the result to analytical choices.
STROBE
The Strengthening the Reporting of Observational Studies in Epidemiology guidelines, a 22-item checklist covering items that should be reported in cohort, case-control, and cross-sectional studies.
Study Identifier
The unique code a registry assigns to identify a specific research study — a clinical trial, systematic review, or other registered research project — as that study itself, distinct from an identifier assigned to one of its resulting outputs (a DOI), its underlying dataset, or an individual researcher (an ORCID iD). A study identifier's format, prefix, and issuing authority are registry-specific: ClinicalTrials.gov assigns 'NCT' numbers, the ISRCTN registry assigns 'ISRCTN' numbers, and PROSPERO assigns 'CRD42' numbers. A single study can carry more than one study identifier if it is registered, or cross-registered, in more than one registry.
Subject repository
A repository the contents of which are connected purely by their discipline, rather than by other factors such as their institutional affiliation (see Institutional Repository)
Technical Metadata
Technical metadata is the category of metadata that documents the digital characteristics of a resource needed to render, verify, and process it correctly over time: file format and version, the hardware and software environment used to create or read it, technical characteristics (size, encoding, resolution, compression), and a record of any technical processing or transformation events (format migration, compression, checksum regeneration) applied since creation. It answers 'what does a machine need to know to open, verify, and correctly interpret this file,' as distinct from what a resource is about (descriptive metadata) or who may use it and under what terms (administrative metadata, in the narrower rights/provenance/access-control sense). Some frameworks, including NISO's influential three-way split, treat technical metadata as a subtype nested under the broader administrative metadata category; others, including the METS schema's own internal structure, give it a distinct named subsection (techMD) alongside rights and provenance subsections rather than folding it into a single undifferentiated administrative bucket. Both framings agree on the underlying content -- format, environment, and fixity information -- they differ only in where they draw the taxonomy line.
Tissue bank
A specific kind of biobank focused on the collection, processing, storage, and distribution of human tissue samples (typically solid tissue specimens from surgical or post-mortem sources), governed under tissue-banking regulation in the relevant jurisdiction.
Trusted digital repository
A digital repository whose mission, governance, technical infrastructure, and procedures have been independently assessed against a recognised standard (e.g. CoreTrustSeal, nestor seal, ISO 16363) and judged trustworthy to preserve digital content over the long term.
UK Data Service (concept)
A UK ESRC-funded data infrastructure that holds, curates, and provides access to social, economic, and population data resources for research, learning, and policy, comprising the UK Data Archive at the University of Essex and partner institutions.
UKRI Funding Service
The UKRI Funding Service is the single online platform (funding-service.ukri.org) that UK Research and Innovation (UKRI) and its seven research councils use to advertise funding opportunities, receive applications, and manage awards through to reporting. An application, account, or award is 'on the Funding Service' when it was created and processed through this platform rather than UKRI's retired Joint Electronic Submission (Je-S) system, which the Funding Service progressively replaced for competitive research council funding from 2023.
Vendor Neutral Archive (VNA)
A vendor neutral archive (VNA) is a medical-imaging storage architecture that ingests, indexes and stores diagnostic images and related clinical content in standard formats (typically DICOM for images, non-DICOM objects via XDS-style wrapping) with vendor-agnostic access interfaces, so the archive is decoupled from any single PACS (Picture Archiving and Communication System) vendor's proprietary database. A storage system counts as a true VNA only if a hospital or health system could, in principle, replace its front-end PACS/viewer without a forced, costly data-migration project — normalization of incoming studies to a consistent format, standards-based query/retrieve (DICOM Query/Retrieve, IHE XDS-I.b, HL7/FHIR for metadata), and no lock-in to a single vendor's schema are the defining tests, not simply 'a central image store.'
VIVO
An open-source semantic-web application and ontology developed by the VIVO community (initially at Cornell University, now under DuraSpace/LYRASIS) that publishes information about researchers, departments, publications, grants, and courses as linked open data and as a navigable web interface.
W3C DCAT (Data Catalog Vocabulary)
DCAT (Data Catalog Vocabulary) is a W3C Recommendation: an RDF vocabulary, published under the namespace http://www.w3.org/ns/dcat#, for describing datasets, dataset series, distributions, and data services within a data catalog so that catalogs published by different organizations can be harvested, aggregated, and searched interoperably. A metadata record qualifies as DCAT (rather than a merely similar, ad hoc catalog schema) when its fields map to DCAT's own RDF classes and properties -- dcat:Catalog, dcat:Dataset, dcat:Distribution, dcat:DataService, dcat:DatasetSeries, and dcat:CatalogRecord, built on top of Dublin Core, FOAF, SKOS, and PROV-O terms rather than a locally invented field set -- and is expressed (or losslessly convertible to) RDF, typically serialized as JSON-LD, Turtle, or RDF/XML.
Watermark Faculty Success (Digital Measures)
A commercial faculty activity reporting (FAR) and faculty-lifecycle management platform, owned by Watermark Insights, used by institutions to collect faculty teaching, research, and service activity data once and repurpose it into annual activity reports, CVs, promotion and tenure dossiers, accreditation reports, and public faculty profiles. It was originally sold as 'Digital Measures Activity Insight' before Watermark's 2018 acquisition and subsequent rebrand to 'Faculty Success.'
Web of Science
Web of Science is Clarivate’s citation-indexing and research-discovery platform, not a single index in its own right. A publication counts as covered by Web of Science only once it has been evaluated and accepted into at least one of the platform’s underlying citation indices — most often the Science Citation Index Expanded, Social Sciences Citation Index, or Arts & Humanities Citation Index within the Web of Science Core Collection, or one of the platform’s regional indices — each governed by its own independent Clarivate selection criteria; being published, or indexed elsewhere such as in Scopus, does not automatically make a title searchable on Web of Science.
What Is a DICOM Conformance Statement?
A DICOM conformance statement is a vendor-published document, structured per DICOM PS3.2 (Conformance), that specifies exactly which SOP (Service-Object Pair) classes a medical-imaging device or software supports and in which network role (SCU or SCP), which transfer syntaxes it can send or receive, which network and media-storage services it offers, and any vendor-specific extensions or restrictions. A document counts as a real conformance statement only if it specifies this level of protocol detail; a marketing claim of 'DICOM compliant' with no published SOP-class/transfer-syntax detail does not meet the bar, because it gives a buyer no way to verify whether two specific devices will actually interoperate.
What Is a Radiology Information System (RIS)?
A radiology information system (RIS) is the software that manages a radiology or imaging department's operational workflow around an exam — patient scheduling and registration, order/accession tracking from request through signed report, radiologist reporting and dictation, results distribution, procedure-based billing/coding, and departmental resource management — as distinct from a PACS, which stores and displays the diagnostic images themselves, and from a hospital's general EHR/HIS, which holds the full patient record across every department, not just imaging.
World Data System certification
Historic certification programme of ICSU's World Data System (WDS) under which scientific data centres in geosciences and related fields were certified as trustworthy; merged with the Data Seal of Approval in 2017 to form CoreTrustSeal.
Zenodo (concept)
A free generalist research repository operated by CERN and developed under OpenAIRE that accepts deposits of datasets, software, publications, presentations, posters, and other research artefacts, minting DataCite DOIs and providing free preservation up to a per-record size limit.

Data management plans (DMPs)

The living documents and machine-actionable expressions that declare how research data will be collected, stored, shared, and preserved.

For the applied guidance behind these terms, see data management plans and funder requirements.

24 terms

Active DMP
A DMP that is actively maintained, updated, and queried during project execution, typically in machine-actionable form, in contrast to a one-off document filed at proposal stage.
Argos (OpenAIRE)
OpenAIRE's open-source DMP authoring service, designed from the outset around the RDA DMP Common Standard and integrated with European Open Science Cloud (EOSC) and OpenAIRE Research Graph services.
Cost element (in DMP)
A line item in a DMP describing a financial commitment associated with data management, such as repository deposit fees, long-term storage, data steward time, or anonymisation services.
Data Management Plan (DMP)
A formal document that describes how research data will be collected, processed, described, stored, shared, preserved, and (where appropriate) destroyed across the lifecycle of a research project.
DataDMP
A DMP authoring and management platform developed in Germany that implements the RDA DMP Common Standard and emphasises integration with institutional research-data infrastructures.
DMP active phase
The phase of the DMP lifecycle during project execution, when the plan is iteratively updated as data are actually generated, processed, and deposited.
DMP assessment
The structured rating of a DMP against a published rubric to produce a comparable score across plans, used in funder evaluation, institutional benchmarking, and capacity-building.
DMP closeout phase
The final phase of the DMP lifecycle, at or after project end, when the plan is reconciled against actual outputs, preservation commitments are confirmed, and the DMP is archived as part of the project record.
DMP compliance check
An automated or rules-based verification that a DMP satisfies the structural and policy requirements of a specific funder, institution, or standard, typically returning a binary or itemised pass/fail outcome.
DMP component
A discrete, reusable section of a DMP corresponding to a logical entity (project, dataset, contributor, host, cost, security_and_privacy) as modelled in the RDA DMP Common Standard.
DMP creation phase
The phase of the DMP lifecycle covering initial drafting, typically at grant-proposal stage, when data types, volumes, and intended sharing arrangements are projected rather than known.
DMP lifecycle
The set of phases through which a Data Management Plan passes from initial drafting at proposal stage, through active project execution, to project closeout and post-project preservation.
DMP narrative
The human-readable prose portion of a Data Management Plan, typically organised under the headings of the applicable funder or institutional template.
DMP review
A formal or informal evaluation of a DMP by a peer, data steward, librarian, or funder reviewer against quality criteria such as completeness, plausibility, and policy alignment.
DMP template
A funder-, institution-, or community-specific structured set of questions and guidance used to elicit the content of a Data Management Plan from researchers.
DMPonline (DCC product)
The UK Digital Curation Centre's hosted instance of the DMPRoadmap platform, providing DMP authoring for UK and international institutions against funder-specific templates.
DMPRoadmap (DMP Tool)
An open-source Ruby on Rails platform for creating, reviewing, and exporting Data Management Plans, jointly developed by the UK Digital Curation Centre and the University of California Curation Center, and deployed under different brands (notably DMPonline and DMPTool).
ezDMP
A funder-focused DMP creation service that guides researchers through funder-specific (originally US NSF directorates') requirements and produces both narrative and structured outputs.
Living DMP
A DMP that is versioned, citable, and intended to evolve over the life of the project, with each significant change captured as a new version of the plan.
Machine-actionable DMP (maDMP)
A Data Management Plan expressed in a structured, machine-readable format (typically JSON conforming to the RDA DMP Common Standard) that enables automated exchange, validation, and updating between systems such as DMP tools, repositories, CRIS/RIMS, and funder portals.
Output management plan (OMP)
A broader successor concept to the DMP that covers all categories of research output (data, software, samples, protocols, models, publications) within a single management plan.
Sharing commitment (in DMP)
A statement in a DMP specifying which datasets will be shared, to whom, on what licence, with what timing relative to project end, and through which infrastructure.
Software management plan (SMP)
A structured plan covering how research software will be developed, documented, licensed, tested, released, and maintained over a project's lifetime, increasingly required alongside or as an extension of a DMP.
Static DMP
A DMP that is produced at a single point in time (typically grant submission) and not subsequently updated, regardless of whether the project's data realities evolve.

Persistent identifiers

The PID ecosystem that disambiguates researchers, organisations, outputs, and projects across every system that touches them.

For the applied guidance behind these terms, see persistent identifiers and identity infrastructure.

40 terms

ARK
Archival Resource Key, a persistent identifier scheme for information objects of any type, in the form ark:/NAAN/Name[Qualifier], where NAAN is a Name Assigning Authority Number and Name is the local identifier; resolvable through any cooperating ARK resolver.
ARK inflection rules
A convention of the ARK identifier scheme whereby appending a single '?' to an ARK URL yields the object's descriptive metadata and appending '??' yields a 'commitment statement' describing the issuing institution's persistence policy for that ARK.
Crossref DOI
A DOI registered through Crossref, the DOI Registration Agency for scholarly publications (journals, books, conference proceedings, preprints, peer reviews, grants), accompanied by metadata deposited in Crossref's XML schema.
Curated org record (ROR)
An entry in the ROR registry that has been reviewed and approved by ROR curators against the published inclusion criteria, carrying a stable ROR ID and metadata including names, types, country, geographic location, parent/child/related/successor/predecessor relationships, and crosswalks.
DataCite consortium
A national or regional grouping of DataCite member organisations led by a 'Consortium Lead' that holds the master agreement with DataCite, allowing member institutions to mint DOIs under a shared fee structure and shared support model.
DataCite DOI
A DOI registered through DataCite, the DOI Registration Agency that serves research data, software, samples, dissertations, instruments, and other non-article research outputs, accompanied by metadata in the DataCite Metadata Schema.
DOI
Digital Object Identifier (ISO 26324), a persistent identifier for an entity (typically a research output) consisting of a prefix assigned to a registrant by a DOI Registration Agency and a suffix assigned by the registrant, resolvable as an HTTPS URI under https://doi.org/.
DOI prefix
The leading portion of a DOI before the first forward slash, of the form 10.NNNN where 10 is the directory indicator for DOI under the Handle System and NNNN is a numeric (or alphanumeric) string assigned by the DOI Registration Agency to a specific registrant.
DOI suffix
The portion of a DOI after the first forward slash, assigned by the registrant within their issued DOI prefix, identifying the specific object; can contain any Unicode characters with the case-insensitivity rule applied during comparison.
DOI tombstone
A tombstone page served at the resolved URL of a DOI after the underlying resource has been withdrawn, providing withdrawal information and metadata while ensuring the DOI itself continues to resolve.
Funder ID
An identifier from the Crossref Funder Registry (formerly FundRef), a curated, open registry of funder names and identifiers used by publishers to tag deposited works with the funders that supported them.
GRID
Global Research Identifier Database, a legacy identifier and registry of research organisations originally operated by Digital Science, frozen to new records in 2021 and superseded by ROR, which seeded its registry from a deduplicated GRID snapshot.
GUID
Globally Unique Identifier, a generic term for an identifier that is intended to be unique across all systems and time, most commonly implemented as a 128-bit UUID but used informally for any opaque, globally scoped identifier.
Handle
An identifier in the CNRI Handle System (RFC 3650-3652), of the form Prefix/Suffix (e.g. 20.500.12345/abcd), resolved by a distributed system of Handle servers that map the identifier to one or more current URLs or other typed data values.
IGSN
International Geo Sample Number, a globally unique persistent identifier for physical samples (geological, environmental, biological) that supports tracking and citation of the sample through subsequent analyses, publications, and derived data.
ISNI
International Standard Name Identifier (ISO 27729), a 16-digit identifier for the public identity of a person or organisation involved in the creation, production, management, or distribution of content, administered by the ISNI International Agency.
ORCID API
The two-tier REST application programming interface (Public API and Member API) operated by ORCID that allows systems to read public ORCID record data and, with researcher authorisation, to read restricted data or write trusted-party assertions to records.
ORCID consortium
A national or regional grouping of ORCID member organisations that share a single membership fee structure and a lead organisation, in order to coordinate ORCID adoption, training, and policy advocacy within a country or region.
ORCID education
An affiliation item in an ORCID record asserting that the iD holder studied at a named organisation, including degree or qualification, department, start and end dates, and the organisation's disambiguated identifier.
ORCID employment
An affiliation item in an ORCID record asserting that the iD holder is or was employed by a named organisation, with start date, optional end date, department, role title, and the organisation's disambiguated identifier (typically a ROR ID).
ORCID iD
A 16-digit persistent identifier, expressed as four hyphen-separated blocks (e.g. 0000-0002-1825-0097) and resolvable as an HTTPS URI under https://orcid.org/, that uniquely identifies an individual researcher across publications, datasets, grants, employments, and peer-review activity.
ORCID record
The structured profile maintained at orcid.org for an individual ORCID iD, containing assertions about the person's names, employments, educations, funding, works, peer reviews, and service activities, each with a visibility setting and a source attribution.
ORCID record permissions
The three-level visibility setting attached to each item in an ORCID record — public, trusted parties only (limited), or private — which the record holder applies individually to names, employments, works, fundings, and other assertions.
ORCID work assertion
A claim, recorded in an ORCID record, that a particular research output (journal article, book chapter, dataset, software, etc.) is associated with the iD holder, with metadata fields including title, type, publication year, external identifiers (DOI, ISBN, PMID), and contributor role.
Persistent URL
An HTTP(S) URL that an issuing organisation commits to maintain unchanged over time so that links continue to resolve correctly even as the underlying resource is moved, renamed, or migrated between systems.
PID consortium
A grouping of PID-provider member organisations, typically at national or regional scale, formed to share infrastructure, contracts, and support around one or more persistent identifier schemes such as DOI, ORCID, or ROR.
PID graph
A graph data structure in which persistent identifiers (ORCID iDs, DOIs, ROR IDs, RAiDs, IGSNs, etc.) are nodes and the metadata relationships among them (creator-of, affiliated-with, funded-by, derived-from) are edges, allowing federated queries across multiple PID-provider registries.
PID minting
The act of generating a new persistent identifier in a registered scheme and registering it, with associated metadata, at the appropriate PID provider so that it becomes resolvable and discoverable.
PID provider
An organisation that issues persistent identifiers from one or more PID schemes, operates (or contracts) the resolution infrastructure for those identifiers, and makes long-term commitments about the maintenance of the identifiers and their metadata.
PID resolution
The process by which a persistent identifier is looked up through its scheme's resolution infrastructure and returned either as an HTTP redirect to the current resource location or as metadata about the resource, depending on the request and the scheme's policy.
PIDINST
A persistent identifier for a research instrument, minted under a DataCite DOI or Handle, conforming to the PIDINST metadata schema developed by an RDA Working Group, that enables citation of and provenance back to the instrument that produced data.
PURL
Persistent Uniform Resource Locator, a URL maintained by a PURL service that redirects (typically via HTTP 302) to the current location of the named resource, allowing the persistent URL to remain stable as the underlying resource location changes.
RAiD
Research Activity Identifier, an ISO-standardised persistent identifier (ISO 23527) for a research project or activity, providing a stable handle around which related people, organisations, outputs, instruments, and funding can be linked over the activity's lifetime.
Resolution service
A networked service that, given a persistent identifier, returns the current location of the named resource (typically by HTTP redirect) or returns its metadata, allowing the identifier itself to remain stable while the resource's location changes.
ROR Curation
The community-driven process by which the Research Organization Registry receives, reviews, and acts on requests to add new organisations, update existing records, merge duplicates, or split records, governed by a published curation policy and managed by ROR's curation team.
ROR ID
A persistent identifier for research organisations issued by the Research Organization Registry (ROR), expressed as an HTTPS URI of the form https://ror.org/0xxxxxxxx where the final nine-character path component is a base32-encoded random value with a check digit.
Tombstone page
A landing page served at a persistent identifier's resolved URL after the underlying resource has been withdrawn, retracted, or made permanently unavailable, providing metadata describing the former resource, the reason for its absence, and (where applicable) a successor identifier.
URN
Uniform Resource Name (RFC 8141), a URI of the form urn:NID:NSS where NID is a registered Namespace Identifier and NSS is the namespace-specific string, intended to denote a resource persistently and independently of any particular resolution mechanism.
UUID
Universally Unique Identifier (RFC 4122 / ISO/IEC 9834-8), a 128-bit value rendered as 32 hexadecimal digits in 8-4-4-4-12 grouping, generated such that the probability of collision across independent generators is negligible.
w3id.org PURL
A persistent URL hosted on the w3id.org domain by the W3C Permanent Identifier Community Group, providing a community-maintained redirect under https://w3id.org/<namespace> to ontologies, vocabularies, and standards documents that may move between hosting providers over time.

FAIR, reproducibility, and sharing

Principles, statements, and practices for making research outputs Findable, Accessible, Interoperable, Reusable, and reproducible.

25 terms

Code availability statement
A statement in a published article describing where the source code used in the study can be obtained, under what licence, and at what version, typically required by journal policy.
Computational environment
The full software and hardware context in which an analysis runs, including operating system, language runtime, library versions, configuration, environment variables, and hardware-specific dependencies (e.g., GPU drivers).
Computational reproducibility
The narrow technical sense of reproducibility: obtaining the same numerical outputs from the same data and code, on a comparable computational environment.
Container image (Docker/Singularity/Apptainer)
A packaged, immutable filesystem and configuration that contains an application together with all its dependencies, runnable identically on any compatible container engine (Docker, Podman, Singularity, Apptainer).
Data availability statement
A statement in a published article describing where the data underlying the study can be found, the conditions of access, and any restrictions, typically required by journal policy.
Data citation principle
Any of the eight principles articulated in the Joint Declaration of Data Citation Principles (Force11, 2014) covering importance, credit and attribution, evidence, unique identification, access, persistence, specificity and verifiability, and interoperability and flexibility of data citations in scholarly communication.
Empirical reproducibility
The ability to obtain consistent observations when an empirical procedure (laboratory, field, or measurement) is independently repeated under matched conditions.
FAIR4RS Software Citation Principles
An extension of the FAIR Guiding Principles to research software, articulating that software should be Findable, Accessible, Interoperable, and Reusable, with the precise interpretations adapted to software's distinctive properties (executability, versioning, dependencies).
Inferential reproducibility
The degree to which independent analysts reach the same qualitative scientific conclusion from the same data, even where their analytical choices differ.
Joint Declaration of Data Citation Principles
The 2014 statement produced by Force11's Data Citation Synthesis Group, signed by a wide community of publishers, funders, repositories, and infrastructure providers, that articulates eight principles for the citation of research data in scholarly communication.
Methods reproducibility
The degree to which a study's methods are reported in sufficient detail that another investigator could re-implement them, independent of whether the same numerical or empirical results would follow.
Open code
The practice of releasing the source code used in a study, under an open-source licence, alongside the publication, such that any reader may inspect, reuse, and re-execute the analysis.
Open data
The practice of making research data freely available for any user to access, use, modify, and share, subject only to attribution requirements, typically through deposit in a public repository under an open licence.
Open materials
The release of the non-data, non-code materials used in a study (stimuli, survey instruments, experimental protocols, training materials, intervention manuals) such that future investigators can re-implement the procedure.
Pre-analysis plan
A detailed, time-stamped document specifying the statistical models, variable transformations, exclusion criteria, and inference rules to be applied to a dataset, lodged before the analyst sees the outcome data.
Pre-registration
The practice of publicly recording a study's hypotheses, design, sample, and analysis plan in a time-stamped registry before data collection or (in secondary-data work) before data access, in order to distinguish pre-specified from post-hoc analyses.
Replicability
The ability to obtain consistent results when an independent investigator collects new data using the same study design and analysis procedures.
Reproducibility
The ability to obtain consistent computational or analytical results when the same data and analysis procedures are applied by an independent investigator using the same code and tools.
Reproducibility audit
A systematic, post-publication examination of whether a study's published results can be obtained from its deposited data and code, typically performed by an independent analyst.
Reproducibility crisis
The widely reported finding that substantial proportions of published research, particularly in biomedical, psychological, and social sciences, fail to reproduce or replicate when re-tested.
Reproducible Research Practices (RRP)
The set of disciplinary norms, tools, and habits that together raise the probability that published research will be reproducible: literate programming, version control, dependency pinning, data deposit, code release, and reporting standards.
Results reproducibility
The narrow sense in which a study's reported quantitative results can be recreated from the deposited data using the deposited analysis procedures.
Software citation (Software Citation Working Group)
The practice of citing research software in the reference list of a publication, with sufficient metadata (authors, title, version, persistent identifier, role) to credit creators and enable retrieval of the cited version.
TOP Guidelines
The Transparency and Openness Promotion Guidelines, an eight-standard framework for journal policies covering citation, data, materials, code, design, analysis, pre-registration, and replication.
Workflow language (CWL/WDL)
A declarative specification language for describing multi-step computational analyses such that the steps, their inputs and outputs, and their software dependencies are portable across compatible workflow execution engines.

This page is the canonical home of the CASRAI RDM glossary, referenced by controlled-vocabulary services including the UN FAO AGROVOC thesaurus. All definitions are CC-BY 4.0.

Frequently asked

RDM Glossary FAQ

What is a research data management glossary?

A research data management (RDM) glossary is a curated, defined set of the terms used to describe how research data is planned for, collected, documented, stored, shared, preserved, and cited across the research lifecycle. The CASRAI RDM Glossary draws 336 of these terms from the CASRAI Dictionary, covering data infrastructure, machine-actionable DMPs, persistent identifiers, research-information systems, and reproducibility.

Who maintains the CASRAI RDM Glossary?

The glossary is maintained by CASRAI as part of the broader CASRAI Dictionary, stewarded by community working groups that draft, review, and ratify entries on a rolling versioned release cadence. The RDM terms are stewarded primarily by the data-infrastructure, machine-actionable DMP, persistent-identifier, research-information-systems, and reproducibility working groups.

Is the RDM Glossary free to use under CC-BY?

Yes. Every entry in the CASRAI RDM Glossary is published under the Creative Commons Attribution 4.0 International licence (CC-BY 4.0), with no paywall and no registration. You may reuse, redistribute, translate, and bundle the definitions commercially or non-commercially, provided you give appropriate credit to CASRAI.

How do I cite a term from the RDM Glossary?

Each term has a stable URI at https://casrai.org/dictionary/term/<slug> together with the dictionary release version. The term page emits a citation widget producing APA, BibTeX, RIS, and Chicago forms. To cite the dictionary release as a whole, use the Zenodo DOI listed on /dictionary/cite.

How does this glossary relate to FAIR and DMPs?

The FAIR principles (Findable, Accessible, Interoperable, Reusable) and data management plans (DMPs) are the operational backbone of research data management, so they are central to this glossary. FAIR-related terms (data citation, persistent identifiers, repositories) and DMP terms (machine-actionable DMPs) each form a dedicated sub-theme below.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →