Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us
Dictionary termTrack BProposedv2026.1

Metadata

Structured information that describes, identifies, or explains a research output — most often a dataset — independent of the content itself, so that the output can be found, correctly interpreted, cited, and reused without the data (or document, sample, or instrument) having to be opened or examined directly. In research data management, metadata is what a repository, catalog, or search index actually reads to determine whether a record matches a query; the underlying data file is opaque to that process until metadata points to it.

ByCASRAI Editorial Board
· Last updated 23 Jul 2026

Examples

Worked examples

  • Is an instance

    A dataset deposited in a repository carries a Title, Creator list with ORCID iDs, PublicationYear, a DataCite DOI as Identifier, ResourceType, Subject keywords, and a Rights/license statement as its metadata record — none of which is data about the phenomenon studied, all of which is data about the dataset.

  • Is an instance

    A field-collected soil-sample dataset includes, alongside the measurement values themselves, metadata recording the collection instrument and calibration date, GPS coordinates, sampling protocol version, and the column definitions and units in the accompanying data dictionary — the information a second researcher needs to interpret and reuse the raw values correctly.

Counter-examples

Looks similar, but isn't

  • Not an instance

    The numeric values inside a spreadsheet cell (e.g. '4.7') are not metadata — they are the data. The column header identifying the variable, its unit of measurement, and a note on how it was captured are the metadata describing that value; the distinction is what the information is ABOUT, not its format or file location.

  • Not an instance

    A folder of raw instrument output files with no accompanying README, data dictionary, or repository record is not 'undocumented metadata' — it has effectively no metadata at all, which is precisely why such deposits routinely fail discoverability and reuse even when the underlying data is sound.

Editorial commentary

Metadata is structured information that describes, identifies, or explains a research output — most often a dataset — independent of the content itself. It is what makes a dataset findable in a repository search, correctly interpretable by a second researcher, and citable as a distinct, identifiable object. In research data management (RDM), metadata is not a bureaucratic add-on to the ‘real’ work of collecting or analyzing data — it is the layer that determines whether that work can ever be found, understood, or reused by anyone other than the person who produced it, including a future version of themselves.

Why metadata matters across the research data lifecycle

Metadata does different work at different stages of a dataset’s life. At the point of collection, it captures provenance — instrument, protocol, collection date, geographic location, personnel — that cannot be reliably reconstructed later. At deposit, it is what a repository indexes and exposes to search engines and aggregators like persistent identifier resolvers, DataCite Search, and OpenAIRE. At reuse, it is the only information a second researcher has before deciding whether a dataset is even relevant to their question, since opening and inspecting the raw file directly is rarely practical at scale. Poor or missing metadata is one of the most common, and most preventable, reasons an otherwise well-collected dataset goes unfound and unreused.

The four commonly recognized types

RDM practice generally distinguishes four functional types of metadata, each answering a different question about a research output:

  • Descriptive metadata — what is this? (title, creator, subject, abstract, keywords)
  • Structural metadata — how are its parts organized? (file relationships, table/column structure, version sequence)
  • Administrative metadata — how was it created and managed? (rights, provenance, funding, access conditions)
  • Preservation metadata — what is needed to keep it usable long-term? (fixity checksums, format migration history)

See Types of Metadata: Descriptive, Structural, Administrative, and Preservation for a full worked breakdown of each, and the dedicated terms for administrative metadata and technical metadata individually.

Metadata standards and schemas

Metadata is only useful for discovery and interoperability if it follows a shared structure other systems can parse. The most widely referenced generic standard is Dublin Core, a fifteen-element set (Title, Creator, Subject, Description, Publisher, Contributor, Date, Type, Format, Identifier, Source, Language, Relation, Coverage, Rights) standardized in ISO 15836:2009, ANSI/NISO Z39.85-2012, and IETF RFC 5013. For datasets specifically minted a DOI, the DataCite Metadata Schema defines six mandatory properties a repository cannot register a dataset DOI without — Identifier, Creator, Title, Publisher, PublicationYear, and ResourceType — plus a recommended tier including Subject, Contributor, Date, RelatedIdentifier, Description, and GeoLocation. Many disciplines layer a domain-specific metadata schema on top of these generic layers to capture field-specific detail a generic standard cannot (e.g. sequencing platform in genomics, survey instrument in social science). See How to Choose a Metadata Schema for a Dataset and a fully worked record at Dublin Core Metadata Record: A Fully Worked Example. The minimal baseline set every record needs regardless of discipline is covered separately as core metadata.

Metadata’s role in FAIR

Metadata is central to two of the four FAIR Guiding Principles: F1 requires that ‘(meta)data are assigned a globally unique and persistent identifier,’ and F2 requires that ‘data are described with rich metadata,’ elaborated by R1 as data being ‘richly described with a plurality of accurate and relevant attributes.’ In practice, a dataset with a persistent identifier but thin or absent metadata is not meaningfully FAIR-compliant — the identifier resolves to something a searcher still cannot evaluate or interpret without opening the file. Metadata is also what makes a dataset’s presence in a Data Management Plan concrete and checkable rather than aspirational: a DMP that commits to ‘documenting the data’ is only verifiable once the actual metadata schema and required fields are specified.

Related terms

See also core metadata, metadata enrichment, metadata schema, administrative metadata, technical metadata, and persistent identifier (PID).

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Metadata"
      vocab-term-identifier="https://casrai.org/dictionary/term/metadata" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/metadata",
  "name": "Metadata",
  "identifier": "https://casrai.org/dictionary/term/metadata",
  "description": "Structured information that describes, identifies, or explains a research output — most often a dataset — independent of the content itself, so that the output can be found, correctly interpreted, cited, and reused without the data (or document, sample, or instrument) having to be opened or examined directly. In research data management, metadata is what a repository, catalog, or search index actually reads to determine whether a record matches a query; the underlying data file is opaque to that process until metadata points to it.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
  "url": "https://casrai.org/dictionary/term/metadata",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "dateModified": "2026-07-23T03:52:21",
  "inLanguage": "en"
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →