Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Scholix: The Framework for Linking Publications and Datasets

Scholix is the RDA/WDS-originated interoperability framework that lets DataCite, Crossref, and OpenAIRE exchange machine-readable links between publications and datasets. This guide explains the information model, governance, and what CRIS managers and repository operators need to do to be part of it.

Scholix (the Scholarly Link Exchange) is an interoperability framework, not a database or a search tool in its own right. It defines a common way for the organizations that already track connections between journal articles and research datasets — DataCite, Crossref, and OpenAIRE among them — to exchange that link information with each other and expose it through a shared query interface. If a repository registers a dataset with DataCite and cites the article it supports, and a publisher registers that article with Crossref and cites the dataset back, Scholix is the agreed format and exchange model that lets those two independently-recorded facts be reconciled into one machine-readable link, rather than two disconnected records sitting in two separate systems.

What Scholix is — and what it isn’t

Scholix is best understood as three things working together, not one product: a consensus among publishers, data centres, and aggregators to exchange data-literature link information systematically; an information model that defines, conceptually, what counts as a “Scholix link” (a source object, a target object, a relationship type, and the organization asserting the link); and a link metadata schema that gives that model a concrete, machine-readable representation. It is deliberately a “wholesaler to wholesaler” framework: individual repositories and publishers don’t exchange links with each other directly under Scholix. Instead, a small number of large hubs — principally DataCite, Crossref, and OpenAIRE — each aggregate links from their own natural communities (data centres registering DataCite DOIs, publishers registering content with Crossref) and exchange those aggregated link sets with each other using the Scholix format.

What Scholix is not: it is not a registry that assigns identifiers, not a peer to DOIs or DataCite DOIs themselves, and not a single centralized database that a CRIS could simply subscribe to. It relies entirely on the identifier and metadata infrastructure that DataCite, Crossref, and other DOI registration agencies already operate, and on those organizations continuing to harvest, normalize, and republish link information through it.

Who created and maintains Scholix

Scholix originated as a joint effort between the Research Data Alliance (RDA) and the International Council for Science’s World Data System (ICSU-WDS), formalized through the RDA/WDS Scholarly Link Exchange (Scholix) Working Group. The framework was publicly launched on 20 June 2016 by RDA and ICSU-WDS together with other data-literature link infrastructure providers. Its formal RDA output — the Scholix metadata schema for exchange of scholarly communication links — went through RDA’s standard community review process and was endorsed in April 2020, after which the working group moved into RDA’s “maintaining deliverables” status rather than continuing as an active drafting group. That matters for anyone evaluating Scholix today: it is a completed, endorsed community specification maintained by its adopting infrastructure providers, not an actively-evolving standard with a visible ongoing committee cycle.

In practice, day-to-day operation of Scholix now sits with the organizations that implement it as data hubs. DataCite and Crossref each expose Scholix-formatted or Scholix-aligned link data drawn from their own registration metadata, and OpenAIRE operates the best-known public aggregation and query layer — the Scholexplorer service — which harvests Scholix links from the participating hubs and makes them queryable through a REST API. Repository operators without a direct route into DataCite or Crossref link infrastructure are generally directed to register their article-dataset links with OpenAIRE’s Scholexplorer as a data source instead.

How the Scholix information model works

A single Scholix link record describes a directional relationship between two scholarly objects, typically a publication and a dataset (though the model doesn’t restrict the object types). Each record carries, at minimum: a source object (identifier, identifier type, and object type — e.g. “literature” or “dataset”); a target object described the same way; a relationship type describing how the source relates to the target; and a link provider — the organization asserting the relationship exists, which carries its own accountability for the claim. Because DataCite’s own metadata schema was deliberately aligned with the Scholix schema, a citation or reference relationship recorded once in DataCite’s relatedIdentifiers/relationType field can be exposed outward as a standards-conformant Scholix link without a separate re-modelling step.

Relationship types draw on a shared, DataCite-aligned vocabulary rather than free text, so the same handful of terms recur across the ecosystem:

Relationship type What it asserts
IsSupplementTo / IsSupplementedBy The dataset is supplementary material for the publication, or vice versa — the most common relationship for data explicitly produced to support one article.
References / IsReferencedBy The publication cites the dataset (or the dataset’s record references the publication) as a source used, without the tighter “supplement” relationship.
Cites / IsCitedBy A more general citation relationship, used where the stricter reference/supplement distinction doesn’t apply.

Because the same underlying relationship can legitimately be asserted from either the publication side (by a publisher registering with Crossref) or the dataset side (by a repository registering with DataCite), Scholix aggregation also has to reconcile directionally-inverse but equivalent claims — part of why the hub model, rather than a flat crawl of every individual repository, is how the framework is actually implemented.

Where Scholix links actually live and how to query them

There is no single “Scholix website” that functions as a search engine for end users. The practical access points are the hubs and aggregators that implement the framework: OpenAIRE’s Scholexplorer service provides a public REST API over harvested Scholix links (drawn from DataCite, Crossref, and other participating sources, and itself part of the wider OpenAIRE research graph); DataCite exposes citation and reference relationships between DOIs through its own APIs, described in its documentation on connecting to, contributing, and consuming citations and references; and Crossref surfaces comparable relationship metadata through its own Crossref Event Data service and article-level metadata, which record events and relationships involving Crossref DOIs, including links to datasets.

For a CRIS manager or repository operator, this means “getting into Scholix” is not a separate registration step so much as a consequence of doing PID metadata well upstream: registering dataset DOIs through DataCite (or a DataCite consortium member) with correctly populated relatedIdentifiers and relationType fields, or, for repositories that don’t mint DataCite DOIs at all, exporting Dublin Core or Scholix-formatted records and registering as a data source with OpenAIRE’s Scholexplorer.

Why Scholix matters for CRIS managers and repository operators

Scholix link data feeds three things research administrators actually care about, beyond the underlying identifier infrastructure itself:

  • Discoverability and reuse. A correctly-asserted IsSupplementTo or References link is what lets a reader (or an automated harvester feeding a CRIS) find the dataset behind a published finding, or the publications that have used a given dataset, without relying on prose mentions in an abstract.
  • Credit and evaluation. Because link assertions are attributable to a named link provider and traceable back to a specific DOI relationship, Scholix-derived data supports the kind of transparent data-citation counting that underpins data-reuse metrics in tenure files, funder reports, and institutional repository statistics — provided the underlying DOI metadata was populated correctly in the first place.
  • Institutional reporting completeness. A CRIS that only ingests publication metadata from Crossref and dataset metadata from DataCite as two unrelated feeds will miss the connective relationship between them unless it also ingests (or queries) the Scholix-aligned relationship data. For institutions reporting on data outputs alongside publications, that gap directly understates data reuse and data-sharing compliance.

The practical failure mode institutions run into is not a Scholix problem as such — it’s upstream metadata quality. If a repository never populates relatedIdentifiers/relationType on its DataCite DOIs, or a publisher never asserts the reciprocal relationship through Crossref, there is no link for Scholix’s hubs to aggregate, regardless of how well the exchange framework itself works.

Frequently asked questions

Is Scholix a database I can search directly?

Not on its own. Scholix is a format and exchange model. The nearest thing to a public search/query layer is OpenAIRE’s Scholexplorer API, which aggregates Scholix-formatted links harvested from DataCite, Crossref, and other participating hubs.

Who governs Scholix today?

Scholix originated as a joint RDA/ICSU-WDS working group output, launched in 2016 and formally endorsed by RDA in April 2020. The working group is now in RDA’s “maintaining deliverables” status; ongoing operation happens through the infrastructure providers that implement the framework, principally DataCite, Crossref, and OpenAIRE.

How is Scholix different from DataCite DOIs or Crossref Event Data?

DataCite and Crossref are DOI registration agencies that assign and hold identifier metadata, including relationship fields between a work and related works. Scholix is the shared format those organizations (among others) use to exchange and expose that relationship data with each other in a common structure, rather than a competing identifier system.

How do I get my repository’s article-dataset links into Scholix?

If your datasets already have DataCite DOIs, populate the relatedIdentifiers/relationType fields correctly (for example, IsSupplementTo pointing at the related article’s DOI). If your repository does not mint DataCite DOIs, you can export Dublin Core or Scholix-formatted metadata and register as a data source with OpenAIRE’s Scholexplorer service directly.

What relationship types does Scholix use?

Scholix reuses the DataCite-aligned relation-type vocabulary, most commonly IsSupplementTo/IsSupplementedBy, References/IsReferencedBy, and Cites/IsCitedBy, applied directionally between a source and target object.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →