Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Metadata Search Engines for Scholarly Research: OpenAlex, Dimensions, Google Scholar, CORE, BASE, and Semantic Scholar Compared

What a scholarly metadata search engine is, how it differs from generic web search and SEO metadata tools, and how OpenAlex, Dimensions, Google Scholar, CORE, BASE, and Semantic Scholar compare for research and reporting tasks.

A metadata search engine, in the research-administration sense, is a tool that indexes the structured, machine-readable metadata describing scholarly works — titles, authors, affiliations, funders, citations, persistent identifiers, and subject classifications — rather than (or in addition to) the full text of those works. Searching one returns a bibliographic record you can filter, export, and link back to a DOI, ORCID iD, or ROR ID, not just a webpage.

This distinguishes the category from two things it’s easy to confuse it with: (1) generic web search engines, which return arbitrary pages ranked by relevance and popularity signals rather than structured scholarly metadata; and (2) SEO metadata tools, which help website owners audit their own <meta> tags and schema.org markup for search visibility — an unrelated, marketing-oriented use of the same word “metadata.” Nothing on this page is about SEO tooling. If that’s what you’re looking for, this isn’t the right page.

What Makes a Search Tool a “Scholarly Metadata Search Engine”

Operationally, a tool belongs in this category if it:

  • Harvests or aggregates bibliographic metadata from publishers, repositories, or other databases — usually via an OAI-PMH harvest, a publisher/Crossref/DataCite feed, or web crawling — rather than hosting the underlying research itself.
  • Structures records around scholarly entities: works, authors, institutions, funders, sources (journals/repositories), and often grants, patents, or clinical trials.
  • Resolves to persistent identifiers where possible — DOIs for works, ORCID iDs for people, ROR IDs for institutions — so records can be deduplicated and linked across systems.
  • Exposes the metadata programmatically, typically via a REST API or bulk data download, not just a search box, since institutional and funder systems need to pull this data into reports, CRIS platforms, and analytics.

The tools below all meet that description but differ sharply in what they index, how open the data is, and what they’re best used for. For related infrastructure concepts, see CASRAI’s guides to persistent identifiers and research information and the Data Management Plan term for how metadata discovery connects to funder compliance.

The Major Scholarly Metadata Search Engines

OpenAlex

OpenAlex is a free, fully open (CC0-licensed) scholarly metadata catalog built and maintained by OurResearch, a nonprofit also known for Unpaywall. It launched January 1, 2022 as a direct successor to Microsoft’s discontinued Microsoft Academic Graph, and structures its index around seven core entities — Works, Authors, Sources, Institutions, Topics, Publishers, and Funders — accessible through a documented REST API and full bulk-data snapshots. Its dataset is CC0 and downloadable in full; as of February 2026 the hosted API itself moved to usage-based pricing for higher-volume calls, while single-record lookups by ID or DOI remain free. See CASRAI’s full guide: OpenAlex: What It Is and How It Works, and the API-specific OpenAlex API guide for endpoint and pagination detail.

Dimensions

Dimensions is a research analytics and metadata database built by Digital Science, linking publications to grants, patents, clinical trials, policy documents, and datasets in a single graph — a broader entity scope than most citation databases, which is why it’s often chosen specifically for funding and policy-impact analysis rather than pure literature search. Access is primarily subscription-based, though Digital Science also offers a free, limited version. See the CASRAI dictionary term: Dimensions (Research Database).

Google Scholar

Google Scholar is the most widely used free scholarly search engine by volume of use, but it is architecturally different from the tools above: it has never offered an official, documented public API, publishes no formal coverage statistics, and its ranking algorithm is undisclosed. It’s a strong first-pass discovery tool for a researcher, but a poor fit for programmatic metadata extraction or reproducible bibliometric reporting — institutions doing that work typically use OpenAlex, Dimensions, Scopus, or Web of Science instead. See Google Scholar API and Google Scholar Citations Profile for what is and isn’t programmatically available.

CORE

CORE (from “COnnecting REpositories”) is a UK-based open-access aggregator, operated by the Open University and Jisc, that harvests metadata and, where permitted, full text from open-access repositories and journals worldwide. It’s built specifically around open-access discovery and full-text access rather than comprehensive scholarly coverage — it indexes on the order of tens of millions of open-access records, drawn from thousands of repositories via OAI-PMH harvesting, and offers a public API and bulk datasets aimed at text-mining and open-access research use cases.

BASE (Bielefeld Academic Search Engine)

BASE is a multidisciplinary academic search engine operated by Bielefeld University Library in Germany. Like CORE, it works by harvesting OAI-PMH metadata — in BASE’s case from a very large number of institutional repositories, subject repositories, and other academic sources (reported figures put it in the hundreds of millions of records from well over ten thousand content providers). A large share of indexed records link to freely accessible full text. BASE is a metadata index first: it doesn’t host content itself, only points to and describes it.

Semantic Scholar

Semantic Scholar is a free, AI-powered academic search engine built by Ai2 (the Allen Institute for AI), publicly launched in November 2015. Beyond standard bibliographic metadata, it layers in AI-generated features — TL;DR summaries, influential-citation classification, and a documented Academic Graph API — that make it useful for literature triage and citation-graph exploration, not just lookup. See CASRAI’s guide: Semantic Scholar: What It Is and How It Works.

How the Major Tools Compare

Tool Data model / scope Access Programmatic access Best suited for
OpenAlex Works, authors, institutions, sources, topics, publishers, funders Free, CC0 data Documented REST API + bulk snapshot Open, reproducible bibliometrics and full-corpus analysis
Dimensions Publications + grants, patents, clinical trials, policy docs Subscription (limited free tier) API (paid tiers) Funding/policy-impact and cross-output analysis
Google Scholar Broad web-crawled scholarly content, undisclosed coverage rules Free None official Quick, informal literature discovery
CORE OA repository/journal metadata + full text where permitted Free Public API, bulk datasets Open-access discovery, text mining
BASE OAI-PMH-harvested repository metadata, very broad source count Free Limited; primarily search interface Broad repository-level OA discovery
Semantic Scholar Works + AI-derived summaries, citation classification Free Documented Academic Graph API Citation-graph exploration, literature triage

For a deeper look at how OpenAlex specifically compares to the two dominant subscription indexes, see CASRAI’s comparison: Scopus vs Web of Science vs OpenAlex.

Choosing the Right One for a Given Task

  • Reproducible bibliometric analysis or an institutional repository crosswalk: OpenAlex — open data, no licensing restriction on reuse, full bulk download.
  • Tracing a publication back to its funding, patents, or clinical trials: Dimensions.
  • A quick, informal literature check: Google Scholar — fast and broad, but not suitable for a report that needs reproducible, programmatically verifiable figures.
  • Finding legally reusable open-access full text at scale: CORE or BASE, both purpose-built around OA repository harvesting.
  • Mapping a citation network or triaging a large literature set quickly: Semantic Scholar’s AI-assisted features.

Research administrators supporting a literature review or evidence synthesis may also want CASRAI’s guide on how to search for peer-reviewed articles in academic databases, and citation-mapping tools such as Litmaps, Connected Papers, and Research Rabbit, which build on top of these same underlying metadata sources rather than replacing them.

How These Tools Relate to Institutional CRIS and Reporting Systems

None of the tools above are a substitute for an institution’s own Current Research Information System (CRIS) or research-information-management platform — they’re external discovery layers an institution queries or imports from, not a system of record for internal reporting. A CRIS typically ingests metadata from one or more of these sources (often via DOI, ORCID, or ROR matching) to populate faculty activity records, grant reporting, and REF/RAE-style assessment exercises. See CASRAI’s broader orientation to that layer of infrastructure on the Persistent Identifiers & Research Information pillar page.

Frequently Asked Questions

Is Google Scholar a metadata search engine?

Only loosely. It returns scholarly results and lets you export a citation, but it has no documented public API, publishes no formal coverage or methodology statement, and isn’t built for programmatic metadata extraction the way OpenAlex, Dimensions, or Semantic Scholar are.

What’s the difference between CORE and BASE?

Both harvest OAI-PMH metadata from open-access repositories and both are free, but they’re run by different organizations (CORE by the Open University/Jisc in the UK; BASE by Bielefeld University Library in Germany), index a different mix of sources, and offer different levels of API/bulk-data access — CORE leans further into full-text and text-mining use cases.

Is OpenAlex free to use?

The underlying data is free and CC0-licensed, including full bulk snapshots. The hosted REST API introduced usage-based pricing for higher-volume calls in February 2026; single-record lookups by ID or DOI remain free regardless of volume.

Which metadata search engine should a research office use for funder reporting?

It depends what the report needs to show. If it needs to link outputs to grants or patents, Dimensions is usually the closer fit. If it needs open, reproducible, fully exportable bibliometrics with no licensing restriction, OpenAlex is generally preferred.

Do these tools replace a university’s CRIS?

No. They’re external discovery and metadata sources a CRIS can import from — not a system of record for internal faculty activity or compliance reporting.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →