Crossref’s 2026 public data file, released in March 2026, gives the clearest picture yet of how far institutional and funding identifiers have spread through scholarly metadata. The annual snapshot — now covering nearly 180 million DOI records — shows Research Organization Registry (ROR) identifiers up 250% year over year and grant-identifier links reaching 50,000 records. The release landed in the same window as Crossref’s second in-person Metadata Sprint, held in São Paulo, which pushed community metadata-quality work into Latin America for the first time. For research administrators and CRIS teams tracking PID coverage, both developments are worth watching together: one is a measurement of adoption, the other is part of the mechanism driving it.
What the 2026 public data file actually contains
Crossref publishes a full data file once a year: every DOI it has registered, with associated metadata, exported in JSON Lines format under a CC0 waiver. The 2026 edition, published March 17, 2026, contains close to 180 million records — 12.7 million of them new since the 2025 file, a 7.6% year-over-year increase. The file is a standard reference point for bibliometricians, PID-infrastructure researchers, and anyone auditing how completely publishers are depositing structured metadata rather than bare citation strings.
What makes the 2026 file notable isn’t just its size. Crossref’s own summary highlights that the records themselves are getting richer — more affiliation identifiers, more funder identifiers, more links between a publication and the grant that funded it — which is a different and arguably more consequential trend than raw DOI volume for anyone relying on this data to answer “who funded this, and where were the authors based” at scale.
ROR adoption up 250%
The headline metadata-quality figure in the 2026 file is a 250% year-over-year increase in ROR identifiers attached to organizations in Crossref records. ROR (Research Organization Registry) is the open, community-governed registry of institutional identifiers built by a collaboration that originally included Crossref, DataCite, and the California Digital Library — see CASRAI’s ROR ID registration guide for how an institution gets and maintains one. Crossref added ROR support to its deposit schema and REST API in 2021, but adoption was gradual for several years; the 2026 figure is the clearest sign yet that publisher-side deposit practices are catching up to the standard, likely helped by Crossref’s parallel move (documented on its blog as “A ROR-some update to our API”) to let ROR IDs be used as funder identifiers, not just affiliation identifiers, anywhere a Funder ID could previously go.
That dual role matters operationally. Before this update, a research administrator reconciling funder data across Crossref and DataCite records often had to hold two different identifier types — Crossref Funder ID for grants metadata and ROR for affiliations — in the same workflow. Consolidating on ROR for both use cases is the kind of quiet interoperability change that shows up in adoption statistics like this one before it shows up in any individual institution’s day-to-day workflow. CASRAI’s ROR/ISNI/Ringgold/GRID crosswalk guide and DOI/ORCID/ROR crosswalk guide cover how these identifiers map to one another in practice, which is useful context for interpreting what a “250% increase” actually changes for anyone reconciling records across systems.
Grant identifiers and the expansion of Grant DOIs
The second figure worth tracking is smaller in absolute terms but arguably more structurally significant: the 2026 file recorded 50,000 records with links to grant identifiers for funding, up from a much smaller base. Crossref has offered a Grant ID / Grant DOI mechanism for several years, allowing a funder to register a persistent identifier for an individual award and letting publishers cite that specific grant — rather than just a funder name and an award number as a free-text string — in the funding metadata of a resulting publication. CASRAI’s Crossref metadata deposit workflow guide covers what a publisher is expected to submit and when, including funding metadata fields.
Grant DOIs solve a real reconciliation problem for research offices and funders alike: without a persistent identifier for the grant itself, connecting “this paper” to “this specific award” at scale depends on exact string matching against award numbers that get typed inconsistently across systems. A growing base of Grant DOI links in Crossref’s data means funders and institutions have more machine-actionable evidence connecting research outputs back to specific awards — relevant for funder compliance reporting, research information systems (CRIS) ingestion, and any institution trying to produce an accurate, non-manual list of “everything this grant funded.”
The São Paulo Metadata Sprint
Alongside the data file, Crossref ran its second Metadata Sprint in São Paulo, Brazil, from March 4–6, 2026 — its first in Latin America, run in partnership with SciELO and conducted across Portuguese, Spanish, and English to support participation from across the region. Roughly 31 participants from Argentina, Brazil, Colombia, Ecuador, and Mexico worked through community metadata-quality problems together, in the same broad tradition as Crossref’s earlier sprint format: bringing publishers, librarians, and metadata specialists into a room to work hands-on on real records rather than treat metadata completeness as a purely automated, back-office problem.
The timing relative to the public data file release is coincidental in the sense that the file is an annual fixed-date export, but it’s not coincidental that Crossref frames both under the same “better metadata, more broadly adopted” push: a regional sprint that builds metadata literacy and deposit practice in a part of the world historically underrepresented in PID adoption statistics is one of the concrete mechanisms behind figures like the 250% ROR increase showing up in the aggregate data a year later.
Why this matters for research administrators and CRIS teams
- Funder reporting: more Grant DOI coverage means fewer manual reconciliations between publication lists and award numbers when producing funder compliance reports.
- Institutional identity: a 250% jump in ROR usage means an institution’s ROR ID is increasingly the identifier publishers and funders already hold for it, reducing the affiliation-disambiguation work a research office has historically had to do by hand.
- CRIS ingestion: richer, more structured Crossref metadata reduces reliance on string-matching heuristics when a current research information system (CRIS) ingests publication records and tries to attach the correct institution and funder.
- Regional equity in PID infrastructure: the São Paulo sprint is a reminder that PID adoption statistics aren’t just a technical curve — they track where community engagement and deposit-practice training have actually reached.
Frequently asked questions
What is Crossref’s public data file?
It’s an annual, complete export of every metadata record for every DOI Crossref has registered, released under a CC0 public-domain waiver in JSON Lines format. The 2026 edition, published March 17, 2026, contains close to 180 million records.
What does “ROR identifiers up 250%” actually measure?
It’s a year-over-year increase in the count of ROR (Research Organization Registry) identifiers attached to organizations — author affiliations and, since Crossref’s API update, funders — within the metadata records in the 2026 file compared to the 2025 file.
What is a Crossref Grant DOI?
A persistent identifier a funder can register for an individual grant or award, distinct from a Funder ID (which identifies the funding organization itself). It lets a publication’s funding metadata point to the specific award that funded it, rather than only naming the funder.
Where was the 2026 Metadata Sprint held?
São Paulo, Brazil, March 4–6, 2026 — Crossref’s second Metadata Sprint and its first in Latin America, run with SciELO and conducted in Portuguese, Spanish, and English.
Does this affect institutions that don’t use Crossref directly?
Indirectly, yes. Crossref metadata feeds PID-graph tools such as DataCite Commons and many CRIS/repository ingestion pipelines, so improvements in Crossref’s own ROR and Grant DOI coverage tend to propagate into the identifier data institutions consume secondhand.
For related PID-ecosystem coverage, see CASRAI’s Persistent identifiers in 2026: ORCID + ROR + RAiD + DOI and Project IDs in 2026: RAiD adoption update, and the comparison of Crossref vs DataCite as DOI registration agencies.







