Metadata, Provenance & Description Standards
Metadata is what makes a dataset findable, interpretable, and trustworthy once it leaves the hands of the people who created it. This sub-cluster covers the standards used to describe research data and digital objects, and the parallel standards used to record provenance: the chain of processes, agents, and inputs that produced a given output. The Dublin Core Metadata Initiative (DCMI) defines a widely adopted core set of descriptive elements (title, creator, subject, date, and so on) that underlies many repository and library systems. The W3C's PROV-O ontology provides a formal vocabulary for representing provenance on the web, allowing systems to express not just what a dataset is but how it came to exist and what transformations it has been through. Domain-specific and infrastructure-specific standards sit alongside these: DataCite's metadata schema and the schema.org Dataset type govern how datasets are registered, discovered, and indexed by search engines and data catalogues, while ISO 19005 (PDF/A) addresses long-term fidelity for archived documents rather than raw data. For a research administrator, metadata standards are not a cataloguing detail but a compliance and discoverability requirement: funders and repositories increasingly specify which schema a submitted dataset must use, and inadequate metadata is a common reason data fails FAIR or repository certification review. Pages in this sub-cluster explain what these schemas and ontologies cover, how they relate to one another, and where a given standard is the appropriate choice for a particular data type or repository.
Guides
DCAT-AP: The EU Application Profile of DCAT for Research Data Portals
DCAT-AP is the SEMIC-maintained EU metadata specification that lets research and open-data portals across Europe be harvested together. This guide covers its mandatory/recommended/optional structure, its GeoDCAT-AP and StatDCAT-AP extensions, and where research data portals actually encounter it.
Frictionless Data / Data Package: A Lightweight Standard for Packaging Tabular Research Data
Frictionless Data’s Data Package standard (Open Knowledge Foundation) packages and validates tabular research data via a lightweight datapackage.json descriptor and Table Schema — how it compares to RO-Crate and DCAT.
Bioschemas: Schema.org Markup Profiles for FAIR-Findable Life-Science Data
Bioschemas defines schema.org profiles for datasets, samples, genes, proteins, and tools so life-science resources become search-engine discoverable and FAIR-findable.
Croissant: MLCommons’ Metadata Format for ML-Ready Datasets
Croissant is MLCommons’ open JSON-LD metadata format, built on schema.org, for describing ML-ready datasets, their features, splits, and labels, so ML tools can load them directly.
RO-Crate: Packaging Research Data and Metadata for FAIR Reuse
RO-Crate is a lightweight, JSON-LD-based specification for packaging research data, workflows, and software with rich, machine-readable metadata. What it is, how it works, and where it fits into FAIR data practice.
Descriptive Metadata for Datasets: Title, Creator, Date, Keywords, and Description
Descriptive metadata is the dataset-level information — title, creator, date, subject keywords, description — that drives discovery and citation, distinct from structural, administrative, and preservation metadata.
Data Dictionary in Research Data Management: Definition, Components, and How It Differs from a Codebook
What a data dictionary is in research data management: its core components (variable names, definitions, units, coded values), a worked example, and how it differs from a codebook and a metadata schema.
How to Cite Data: Formats and the Data Citation Principles
Citing a dataset differs from citing a journal article. This guide covers the core elements of a data citation, DataCite’s recommended format, how APA/MLA/Chicago handle datasets, and the FORCE11 Joint Declaration of Data Citation Principles.
Dublin Core Metadata Record: A Fully Worked Example
A complete, fully filled-out Dublin Core metadata record for an illustrative research dataset, covering all fifteen elements plus Simple vs. Qualified Dublin Core.
Electronic Lab Notebooks for Chemistry: Structures, Reactions, and CAS Registry Linking
What changes when an electronic lab notebook has to be chemistry-aware: structure drawing and editing, reaction-centric data capture, CAS Registry Number and InChI linking, and how to evaluate a chemistry ELN platform.
Electronic Lab Notebook Template: Fields, Structure & Audit Trail
What a well-structured electronic lab notebook entry actually contains: metadata fields, experiment body structure, versioning and audit-trail elements, and e-signature/witness sign-off blocks.
Types of Metadata: Descriptive, Structural, Administrative, and Preservation
The four functional types of metadata (descriptive, structural, administrative, and preservation) explained with research-data examples, and how each relates to a metadata schema.
Data Provenance: Definition, Standards, and How to Document It in Research
What data provenance means, how it differs from lineage and metadata, and the standards (PROV-O, RO-Crate) used to document it in research.
CASRAI’s Picklists and Controlled Vocabularies: How They Work and How to Use Them
A field guide to CASRAI’s own picklist system: the 10 published controlled vocabularies, how they fit into the Dictionary’s object-template model, and how to browse, query, or reuse them.
How to Cite a Dataset: A Practical Guide with Real Examples
What a complete dataset citation needs (creator, title, publisher, year, version, DOI/identifier), with real worked examples in DataCite’s recommended format and APA 7th edition style.
How to Choose a Metadata Schema for a Dataset
A decision framework for choosing between Dublin Core, the DataCite Metadata Schema, DDI, schema.org/Dataset, and discipline-specific schemas based on repository requirements, discoverability goals, and data type.
Making Your Dataset Discoverable in Google Dataset Search
Google Dataset Search indexes schema.org Dataset markup, not dataset descriptions in prose. For most researchers that markup is generated automatically from their repository’s DataCite metadata record — this guide explains the actual crosswalk and what to get right at deposit time.







