OpenAlex has published the first public quantification of a problem that research security offices have been working around informally for years: affiliation strings matched to the wrong institution, including matches onto research security watch lists. In a 27 July 2026 blog post, the OpenAlex team reported that it had cross-referenced entries on major watch lists against its affiliation data and removed 6,191 over-merges — cases where a researcher’s affiliation string had been incorrectly linked to one of 192 listed organizations.
What OpenAlex found
OpenAlex builds its affiliation data by algorithmically matching the free-text institution strings that appear on scholarly works — more than 150 million unique strings, drawn from over 500 million works — to structured organization records in the Research Organization Registry (ROR). For a large, well-known institution this works reliably: the post notes that the University of British Columbia alone has roughly 200,000 distinct text-string variants that all correctly resolve to its single ROR ID. Across most academic institutions, OpenAlex reports better than 98% precision and better than 90% recall for this matching process.
The failure mode sits at the tail of that distribution. Smaller organizations, companies, and institutes with short or generic names — and particularly organizations that have no ROR record of their own — are far more likely to be matched by acronym or partial-name similarity to something else entirely. OpenAlex’s post gives several concrete examples of exactly this happening:
- A small US biotechnology firm, identified in its publications by a short name, was linked to a Russian research institute because the firm had no ROR record and the closest available string match was the Russian institute’s acronym.
- A Japanese engineering company division (“System Equipments Div., Asahi Engineering Co., Ltd.”) was matched to a Bureau of Industry and Security-listed Chinese entity called “System Equipment,” on the basis of shared wording alone.
- A company named “Engineering-Physics, Inc.” was linked to the Russian Institute of Engineering Physics, again through terminology overlap rather than any actual institutional connection.
None of these are edge cases in the sense of being rare mismatches on obscure records; they are the predictable output of matching short, generic English-language business names against a global affiliation-string corpus using text similarity. OpenAlex screened roughly 2.5 million affiliation strings against the watch lists as part of this audit, covering an estimated 4.24 million works, and removed the 6,191 confirmed bad matches it found in that pass.
Why this matters for research security screening
Bibliometric affiliation data has quietly become an input to institutional research security compliance workflows, not just an analytics convenience. Under NSPM-33 and the certification and disclosure regimes that have followed it — including agency-specific requirements from NIH, NSF, DOE, and others — institutions and research security officers increasingly rely on automated screening of author affiliations against lists of organizations of concern to flag potential undisclosed foreign talent-program ties, sanctioned-entity connections, or other export-control-relevant relationships. OpenAlex’s affiliation-to-ROR matching pipeline is one of the data sources that kind of screening can draw on, whether directly or through downstream tools built on top of it.
A false positive in that pipeline is not a cosmetic data-quality issue. Where an automated match flags a researcher as affiliated with a watch-listed organization, the practical consequence for that researcher can be a compliance inquiry, a hold on funding disbursement, or a demand to explain a connection that does not actually exist — before anyone has confirmed the underlying string match was correct. Given that OpenAlex’s own reported precision, while high in aggregate, still leaves millions of potential mismatches across a corpus of this size, and given that the organizations most likely to be mismatched are exactly the small, ROR-record-less ones that watch lists often also contain, the failure mode disproportionately concentrates on precisely the matches a screening process most needs to get right.
What OpenAlex says it is asking for
Rather than presenting the 6,191-correction cleanup as a closed issue, OpenAlex frames it as a starting point. The post explicitly invites research security professionals to work with the project: to report additional mismatches, to help document the matching pipeline’s known limitations honestly, and to contribute to building more auditable infrastructure around affiliation data specifically for this use case. OpenAlex is candid that millions of potential errors likely remain across its 500-million-plus work corpus despite the high aggregate precision rate, and that organizations without a ROR record of their own will remain disproportionately vulnerable to this kind of mismatch until that gap is closed at the source — either by those organizations registering in ROR, or by improvements to the matching model itself.
The practical takeaway for institutions
For research security offices and compliance staff using bibliometric affiliation data — whether pulled directly from OpenAlex, from a vendor tool built on it, or from any other affiliation-matching pipeline — the OpenAlex disclosure is a concrete reminder that an automated watch-list hit is a lead to verify, not a finding to act on. A string match should prompt a manual check of the underlying source document, the researcher’s actual institutional history, and, where a ROR ID is available, confirmation that the ID genuinely resolves to the entity in question, before any inquiry, hold, or disclosure demand reaches the individual researcher. That verification step is precisely what OpenAlex’s own post is asking the research security community to help build durable infrastructure around, rather than leaving it to ad hoc case-by-case correction after the fact.
Frequently asked questions
What did OpenAlex actually change?
OpenAlex cross-referenced its affiliation-matching output against major research-security watch lists and removed 6,191 cases where a researcher’s affiliation string had been incorrectly linked to one of 192 listed organizations, publishing the methodology and examples in a 27 July 2026 blog post.
Does this mean OpenAlex affiliation data is unreliable?
No — OpenAlex reports better than 98% precision and better than 90% recall for affiliation matching across most academic institutions. The error concentration is specific: small organizations, companies, and institutes with short or generic names and no ROR record are disproportionately likely to be mismatched, not the corpus as a whole.
Should institutions stop using bibliometric data for research security screening?
OpenAlex’s own framing is not to abandon the practice but to treat an automated match as a starting point requiring manual verification, and to support building more auditable matching infrastructure rather than relying on the current pipeline as a final determination.







