Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

How to Compile One Complete Publication List Across PubMed, Scopus, Web of Science, and Google Scholar

A practical workflow for pulling and deduplicating publication records across PubMed, Scopus, Web of Science, and Google Scholar into one complete, accurate list for a tenure dossier, biosketch, or institutional report.

No single database indexes everything a researcher has published. PubMed is limited to biomedical/life-science journals; Scopus and Web of Science each apply their own journal-coverage criteria and miss different subsets of conference proceedings, book chapters, and regional journals; Google Scholar indexes the broadest range (including preprints, theses, and grey literature) but has no reliable author-disambiguation layer of its own. A publication list built from only one source—for a tenure dossier, an annual research report, a grant biosketch, or an institutional repository—will typically be incomplete, contain duplicates once merged with anything else, or both. This guide sets out a practical workflow for pulling records from all four sources, matching and deduplicating them, and producing one clean, complete list.

Why one database is never enough

  • PubMed/MEDLINE only indexes biomedical, life-science, and related journals selected under NLM’s own inclusion criteria; it will not surface work published in engineering, social-science, or humanities venues, and it indexes relatively few conference proceedings.
  • Scopus and Web of Science each apply independent, publisher-negotiated journal-coverage lists. They overlap substantially but not completely—a journal indexed in one is not guaranteed to be indexed in the other, and neither covers all conference proceedings, working papers, or non-English-language venues equally.
  • Google Scholar has the broadest coverage of the four—it crawls preprint servers, institutional repositories, theses, and publisher sites directly—but it has no controlled author-identity system of its own. A Google Scholar profile is self-curated by the researcher (or an administrator on their behalf); nothing algorithmically verifies that every listed item belongs to that person, and nothing prevents a real item from being missed.

Because coverage only partially overlaps, compiling one complete list means pulling from all four and reconciling the differences, not picking the single best source.

Step 1: Establish the identifier you’ll search with

Before searching any database, confirm the researcher’s ORCID iD and, where they exist, their database-specific identifiers—Scopus Author ID and Web of Science ResearcherID (now surfaced as part of a researcher’s Web of Science Researcher Profile). These identifiers are the single biggest lever for accuracy: searching by name alone runs into common surname collisions, name-order conventions that vary by culture, middle-initial inconsistency between publications, and institutional affiliation changes over a career—all of which cause both missed records (false negatives) and wrongly attributed records (false positives). See CASRAI’s guides on Web of Science author search and disambiguation and claiming a Web of Science Researcher Profile, and on optimizing an ORCID profile for discoverability, for the mechanics specific to each system.

If no Scopus Author ID or ResearcherID exists yet, note that Scopus IDs are algorithmically generated by Elsevier once a researcher has at least two Scopus-indexed documents—they are not requested or self-registered—and the matching algorithm frequently splits one person’s output across multiple author-ID profiles because of name variants, affiliation changes, or subject-area shifts. If a search turns up more than one profile for the same person, that is expected, not an error; each one needs to be searched and its results merged before deduplication.

Step 2: Pull records from each database

PubMed

Search by author name combined with affiliation filters to narrow ambiguous names (PubMed’s search syntax supports [Author] and [Affiliation] field tags, and an ORCID-linked search where the author has associated their ORCID with their PubMed-indexed publications). Once results are confirmed, use PubMed’s own Send to > File option: select the Summary (text) format, or, for import into a reference manager, the citation-manager export, which produces an .nbib file in MEDLINE format compatible with EndNote and most reference managers. Records can also be pulled programmatically via NCBI’s E-utilities (esearch/efetch) for anyone comfortable scripting a bulk pull, which is worth considering for a large or recurring compilation (e.g., an annual institutional report) rather than a one-off dossier.

Scopus

Use Scopus’s Author Search, ideally starting from a known Scopus Author ID rather than a name search, then check for and merge any split/duplicate author profiles the search turns up (Scopus’s own author-merge request workflow handles this at the source, though it is not required just to compile a list—pulling from all matching profiles manually works too). Export results in CSV, BibTeX, or RIS format; note that Scopus only includes the DOI field in the complete-format export or a custom field-selection export, not in the default citation-only export, so choose a fuller export format if DOI-based deduplication (see Step 3) is the plan.

Web of Science

Use the Web of Science Researcher Profile (built around a ResearcherID) where one exists, or Basic/Advanced Search by author and organization otherwise—see CASRAI’s guide to searching Web of Science effectively for query-syntax specifics. Export in RIS or a similar reference-manager-compatible format for the same reason as Scopus: RIS is the common denominator most reference managers (Zotero, EndNote, Mendeley) can import directly, which matters for the merge step next.

Google Scholar

If the researcher maintains a Google Scholar profile, it is the fastest single pull, but treat it as a starting list to verify rather than a finished one—Google Scholar profiles are self-maintained, so both omissions (a paper the researcher never added) and inclusions that do not actually belong to them (a common problem with common names, or where Scholar’s automated suggestion algorithm mis-attributed a paper) are possible. From a profile, select the relevant items and use the export function to produce a CSV or a citation-manager format; note that Scholar’s own interface caps display/export at a limited number of items per page (adjustable up to 20 in Scholar’s settings, well below what a senior researcher’s full output may run to), so a full pull of a large profile takes several export passes rather than one. There is no official bulk API for Google Scholar; third-party tools exist to script larger pulls but are outside Google’s supported interface and worth flagging as such if used for anything that needs to be defensible (e.g., a formal report).

Step 3: Merge and deduplicate

Import all four exports into a single reference manager (Zotero, EndNote, and Mendeley all accept RIS and BibTeX, which covers every export format above) and deduplicate there rather than by eye. Practical matching strategy, roughly in order of reliability:

  1. DOI match — the most reliable single field when present; most reference managers can auto-detect exact DOI duplicates. This is why pulling a DOI-inclusive export format from Scopus (see Step 2) matters.
  2. Title + author + year match — the fallback for records without a DOI (conference proceedings, some book chapters, older records). Reference-manager dedup tools generally do fuzzy title matching to catch near-identical titles with minor punctuation/capitalization differences between databases.
  3. Manual review of near-matches — automated dedup will flag likely-but-not-certain duplicates; these need a human check, particularly for preprint vs. published-version pairs (the same underlying work, legitimately two records if the compilation is meant to distinguish them, one record if not) and for conference-paper-then-journal-extension pairs, which are related but genuinely distinct outputs.

Decide up front, before merging, how the list should treat: preprints that were later formally published (keep one, the other, or both, clearly labeled); self-authored items appearing in multiple author-ID profiles for the same person (a Scopus-merge or ResearcherID-merge situation from Step 1); and non-English or non-journal outputs that may appear in Google Scholar but not the three indexed databases. These policy decisions, not the mechanical dedup step, are usually where inconsistency creeps into institutional publication lists—write the rule down once and apply it consistently.

Step 4: Reconcile what each source missed

After deduplication, compare the merged list against each source individually to catch what would not show up in a straight merge: a paper Google Scholar indexed that never appeared in any author-ID search on Scopus or Web of Science (common for conference proceedings or non-English venues), or a PubMed record that lacked an ORCID link and so was not caught by an identifier-based search in the other three. This reconciliation pass is why Step 1’s identifiers matter—without them, this step is effectively a second full name-based search across every database, with all the same disambiguation risk as the first pass.

Common pitfalls

  • Name variants going unmatched — a middle initial present in one database’s record and absent in another’s is a common cause of two records for the same paper surviving a dedup pass; check name-format consistency, not just DOI/title, when reviewing near-duplicates.
  • Preprint/published-version confusion — treating a preprint and its later peer-reviewed version as either automatically the same record or automatically two unrelated ones, without a stated rule, produces an inconsistent list across researchers in the same report.
  • Split author profiles left unmerged — searching only one of several Scopus Author ID or ResearcherID profiles for a researcher whose output got algorithmically split will silently under-count their work; check for split profiles before finalizing.
  • Treating Google Scholar as authoritative — because it is self-curated and has no verification layer, an unreviewed Google Scholar export can both miss real items and include mis-attributed ones; always cross-check against at least one indexed, identifier-based source before treating a Scholar-derived list as final.
  • Export-format mismatches — pulling a citation-only export (no DOI, no abstract) from one source and a full-record export from another makes automated dedup less reliable simply because the fields available to match on differ; standardize the export format across sources where the tool allows it.

Frequently asked questions

Which database should I start with?

Start wherever the researcher’s identifiers are most complete—typically ORCID first (since it can be linked to records in PubMed, Scopus, and increasingly Web of Science), then whichever of Scopus or Web of Science has an established, unsplit author profile. Treat Google Scholar as a completeness check pulled last, precisely because it needs the most manual verification.

Do I need a subscription to Scopus or Web of Science to export records?

Export functionality in both requires institutional access (a Scopus or Web of Science subscription, typically via a university or research-institution library). PubMed and Google Scholar are both free to search and export from.

How often should a compiled publication list be refreshed?

For anything used in an ongoing context—a lab website, an institutional CRIS record, an annual report—treat this as a repeatable process rather than a one-time pull, and re-run it at whatever cadence matches the researcher’s publication rate and the list’s use (annually is typical for institutional reporting; before submission is typical for grant biosketches and tenure dossiers).

Can this be automated?

Partially. NCBI’s E-utilities allow scripted PubMed pulls, and both Scopus and Web of Science offer institutional APIs for programmatic access where the subscribing institution has enabled them. Google Scholar has no official API, which is the main reason full automation across all four sources is not currently possible end to end—the Scholar step generally stays manual or semi-manual.

Related CASRAI resources

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →