Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

OpenCitations and I4OC: Open Citation Data vs. Web of Science and Scopus

How OpenCitations and the Initiative for Open Citations (I4OC) provide free, CC0-licensed, machine-readable citation data as an open alternative to subscription citation counts from Web of Science and Scopus.

OpenCitations is an independent, not-for-profit infrastructure organization that publishes scholarly citation data as free, machine-readable, open Linked Data — and the Initiative for Open Citations (I4OC) is the advocacy coalition that persuaded the scholarly-publishing ecosystem to actually make that data available in the first place. Together they form the closest thing research administration currently has to an open alternative to citation counts drawn from subscription databases such as Web of Science and Scopus: citation links that anyone can query, download in bulk, and reuse without a license, because they are released under a CC0 public-domain waiver rather than held behind a paywall.

What I4OC actually did

The Initiative for Open Citations launched in 2017, founded jointly by OpenCitations, the Wikimedia Foundation, PLOS, eLife, DataCite, and the Centre for Culture and Technology at Curtin University. I4OC itself does not host or produce citation data — its role has been advocacy and coordination, persuading publishers who deposit reference lists with Crossref to also mark those references as open, so that anyone (not just Crossref members) can retrieve them via the Crossref API.

That advocacy had a measurable effect. Before I4OC began, only around 1% of journal articles with reference lists deposited in Crossref made those references openly available; by August 2021, I4OC and OpenCitations reported that figure had risen to roughly 88% of the 56.1 million Crossref-indexed articles carrying reference data. Crossref’s own subsequent policy changes pushed the practice further toward becoming the default for new deposits, meaning the raw material feeding OpenCitations and similar open-citation efforts is now far more complete than it was a decade ago. Readers should treat the exact current percentage as a moving target rather than a fixed figure — check i4oc.org or OpenCitations’ own reporting for the latest count.

What OpenCitations is, and who runs it

OpenCitations began in 2010 as a one-year pilot project funded by JISC, directed by David Shotton at the University of Oxford’s Department of Zoology. Silvio Peroni of the University of Bologna later joined as co-director, and OpenCitations is now administered by the Research Centre for Open Scholarly Metadata at the University of Bologna, with Peroni and Shotton as director and associate director respectively. It operates as an independent, not-for-profit infrastructure organization rather than a commercial vendor, and has drawn funding at various points from the Wellcome Trust, the Alfred P. Sloan Foundation, and European Union Horizon 2020/Horizon Europe grants (including the OpenAIRE-Nexus and GraspOS projects).

The data model: RDF, SPAR ontologies, and CC0

OpenCitations publishes its data as Linked Open Data using the OpenCitations Data Model (OCDM), built on the SPAR (Semantic Publishing and Referencing) Ontologies. Citations are treated as first-class data entities in their own right — not just an incidental side-effect of bibliographic metadata — described using the Citation Typing Ontology (CiTO) for citation function and type, and the Provenance Ontology (PROV-O) to record who asserted a given citation, when, and from what source. Every citation link also receives its own OpenCitations Identifier (OCI), a persistent identifier for the citation itself rather than for either endpoint. All data is released under a CC0 public-domain waiver, meaning it can be reused, redistributed, and built into institutional systems without attribution requirements or licensing negotiation — a structural difference from Web of Science and Scopus, whose citation data can only be redistributed or bulk-exported within the terms of a paid subscription.

The OpenCitations Index, COCI, and OpenCitations Meta

The core citation dataset is the OpenCitations Index, a continuously growing collection of DOI-to-DOI (and other persistent-identifier-to-persistent-identifier) citation links. It began as COCI — the OpenCitations Index of Crossref open DOI-to-DOI citations, launched in 2018 and drawn entirely from Crossref’s open reference data — and has since expanded to ingest citation data from additional sources, including DataCite, the NIH Open Citation Collection, OpenAIRE, and the Japan Link Center, growing the Index to several hundred million citation links. A companion dataset, OpenCitations Meta, holds the bibliographic metadata (titles, authors, venues, dates) for the citing and cited works themselves, versioned using the same provenance model. All of it is queryable through a public SPARQL endpoint, a REST API, and downloadable bulk RDF dumps — no institutional login or subscription required.

OpenCitations/I4OC vs. Web of Science and Scopus: what actually differs

These aren’t drop-in substitutes for each other, and treating them as interchangeable in a report or dossier will misrepresent what each one measures. The practical differences that matter for research administration:

  • Access and licensing. Web of Science and Scopus are subscription products with contractual restrictions on bulk export and redistribution. OpenCitations data is CC0 — free to query, download in bulk, and embed in an institutional CRIS or dashboard without a license.
  • Coverage and curation. Web of Science and Scopus apply editorial selection criteria to decide which journals and sources are indexed at all, and both have decades of retrospective, non-DOI-dependent coverage. OpenCitations’ coverage is a function of what publishers have deposited with Crossref (and the other source databases) as DOI-registered, open reference data — which means excellent recent coverage for compliant publishers, but real gaps for older literature, publishers that haven’t opened their references, and content that never received a DOI in the first place.
  • What’s actually being counted. A citation link in OpenCitations is a raw, machine-readable DOI-to-DOI (or PID-to-PID) relationship with type/function metadata attached (via CiTO) where available. Web of Science and Scopus citation counts sit inside proprietary indexing and de-duplication pipelines and feed journal-level indicators (the Journal Impact Factor via Web of Science, CiteScore via Scopus) that OpenCitations does not itself compute.
  • Institutional role. Web of Science and Scopus remain the default sources many funders, tenure committees, and national research-assessment exercises expect to see cited, precisely because of that long editorial track record. OpenCitations is not yet a like-for-like substitute in those specific contexts, but it is increasingly used alongside them — particularly for transparency, reproducible bibliometric research, and feeding open infrastructure such as CRIS platforms and citation graph tools that need openly licensed data to build on.

See Is Web of Science Citation Data Reliable for Research Assessment? for a closer look at what a Web of Science citation count does and doesn’t measure, and Lens.org vs. Scopus for how a different open alternative compares directly against a subscription incumbent.

Why this matters for research administrators

Because OpenCitations data carries no licensing restriction, it can be pulled directly into institutional research information systems, open-science dashboards, or funder-compliance reporting without negotiating redistribution rights — a real, practical advantage over subscription citation databases when the goal is building internal tooling rather than looking up a single citation count. It also complements the rest of the persistent identifier layer that research administrators already work with: OpenCitations links resolve through DOIs, the same identifiers used across DataCite deposits, ROR institutional records, and ORCID author records, so citation data derived from it slots into the same identifier graph rather than requiring a separate crosswalk. For grant reporting or tenure files specifically, most institutions still treat Web of Science or Scopus counts as the expected reference point; OpenCitations is best positioned today as a supplementary, verifiable, openly licensed source — useful for spot-checking a count, building custom analysis, or citing a transparent methodology — rather than a wholesale replacement in contexts where a specific proprietary metric is explicitly required.

Limitations to disclose, not just use

An open data source is not automatically a complete one. OpenCitations’ coverage depends entirely on what upstream sources (Crossref, DataCite, NIH-OCC, OpenAIRE, JaLC) have deposited, so a work published by a non-compliant publisher, indexed only in a database OpenCitations doesn’t yet ingest, or predating the DOI system, may show a citation count that understates its real influence. Any research-assessment use of OpenCitations data should disclose this the same way any single-source citation count should be disclosed — as a measure of citation activity within a specific indexed universe, not a neutral, complete count of scholarly influence.

Frequently asked questions

Is OpenCitations data actually free to use?

Yes. All OpenCitations datasets are released under a CC0 public-domain waiver, with no login, subscription, or attribution requirement, accessible via SPARQL endpoint, REST API, or bulk RDF download.

Is I4OC still active, or was it a one-time campaign?

I4OC’s core advocacy goal — getting publishers to open their Crossref-deposited references — has now been achieved for the large majority of Crossref-indexed content, so its role today is more monitoring and continued coordination than the original persuasion campaign. Check i4oc.org directly for current status and figures.

Can OpenCitations replace Web of Science or Scopus for tenure and grant reporting?

Not as a direct substitute in contexts where a funder or committee specifically expects a Web of Science or Scopus figure — those remain the institutionally expected reference points in many assessment processes. OpenCitations is better used today as a supplementary, openly licensed, verifiable source alongside them.

What is the difference between OpenCitations, COCI, and the OpenCitations Index?

COCI was the original OpenCitations Index dataset, built solely from Crossref’s open reference data. The OpenCitations Index has since grown to aggregate citation links from several source databases beyond Crossref, so COCI is now best understood as the founding, Crossref-derived component of a broader Index rather than a separate product.

How does OpenCitations relate to persistent identifiers like DOI and ORCID?

OpenCitations’ citation links are built on DOI-to-DOI (and more broadly PID-to-PID) relationships, and each citation itself receives its own persistent OpenCitations Identifier (OCI) — so it sits directly on top of, and interoperates with, the same DOI/ORCID/ROR identifier infrastructure used elsewhere in research information management.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →