Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Bibliometric Analysis: Methodology, Workflow, and How to Conduct One

What bibliometric analysis is, the step-by-step workflow for conducting one, and a landscape overview of techniques from citation counts to co-word analysis.

Bibliometric analysis is the quantitative study of publications, citations, and related metadata to describe patterns in a body of scholarly literature — how a field has grown, which works and authors sit at its center, how its sub-topics relate, and how its output is distributed across journals, institutions, or countries. The term itself dates to a 1969 paper by British librarian Alan Pritchard, “Statistical Bibliography or Bibliometrics?”, which defined it as the application of mathematical and statistical methods to books and other media of communication.

This guide covers bibliometric analysis as a research method: what it is for, the workflow used to conduct one, and the landscape of techniques available. It is the upstream overview for two more specific CASRAI guides — Citation Network Analysis, which covers co-citation and bibliographic coupling in depth, and VOSviewer, which covers the most widely used bibliometric-mapping software — so if you already know you need one of those specific techniques or tools, those pages go further than this one does. This page is for readers who need the full picture first: what bibliometric analysis actually involves end to end, and which technique fits which question.

What Bibliometric Analysis Is For

Bibliometric analysis answers structural and evaluative questions about a literature that close reading of individual papers cannot answer at scale: which sub-topics make up a field, how the field’s boundaries and vocabulary have shifted over time, which authors, journals, or institutions are most influential or productive, and how research output and collaboration are distributed geographically or organizationally. It is used for:

  • Mapping a field before or alongside a traditional literature review, to identify major clusters, seminal works, and research fronts a keyword search alone might miss.
  • Research assessment — describing an institution’s, journal’s, or researcher’s output and influence using publication and citation counts, though see the caveats on metric-only assessment below.
  • Science mapping and science policy — tracking how a discipline or funding area has evolved, and where emerging topics are gaining momentum.
  • Journal and publication-venue analysis — comparing scope, output, and citation performance across journals in a domain.

It is a quantitative complement to, not a replacement for, close reading. A widely cited methodological overview by Naveen Donthu and colleagues, published in the Journal of Business Research in 2021, frames bibliometric analysis specifically as a way to handle volumes of literature too large to review manually while still surfacing a field’s structure and evolution — the method scales where narrative review does not, but it cannot substitute for a reviewer’s judgment about the substantive content of any individual paper.

When to Use It (and When Not To)

Bibliometric analysis is well suited to a research question framed around a body of literature as a whole: how has a field evolved, who are its most influential contributors, which sub-topics exist and how do they relate, which journals dominate a domain. It depends on having a large enough, well-defined corpus to analyze — a handful of papers doesn’t produce meaningful patterns, and an overly broad or poorly bounded search returns a corpus too noisy to interpret. It is a poor fit for questions that require evaluating the methodological quality or substantive findings of individual studies (that is the job of a systematic review or meta-analysis) or for very new, small literatures that haven’t yet accumulated enough publications or citations to show a pattern.

The Bibliometric Analysis Workflow

Most published bibliometric analyses, and the methodological guidance describing how to conduct one, follow a broadly consistent sequence of steps.

1. Define the research question and scope

Before touching any database, the analysis needs a bounded question: which field, sub-field, or research question is being mapped, over what time period, and at what unit of analysis (documents, authors, journals, institutions, countries, or keywords/terms). This scoping decision drives every downstream choice, including which search strategy and which bibliometric technique will actually answer the question.

2. Choose the data source(s)

Bibliometric analysis depends entirely on the coverage and metadata quality of the source database, so this choice materially affects the result. The major options:

  • Web of Science (Clarivate) — a long-running, curated citation index with deep historical coverage, widely used as the traditional standard for bibliometric work, particularly in the sciences.
  • Scopus (Elsevier) — broader journal coverage than Web of Science in many fields, especially outside North America and Western Europe, and a common alternative or complement.
  • Dimensions (Digital Science) — links publications to grants, patents, clinical trials, and policy documents alongside citations, useful when the analysis needs to connect research output to funding or downstream impact.
  • OpenAlex and other open bibliographic sources — free, openly licensed alternatives to the subscription indexes above, with rapidly growing coverage; see CASRAI’s guide to OpenAlex use cases and the broader comparison of metadata search engines for scholarly research (OpenAlex, Dimensions, Google Scholar, CORE, BASE, and Semantic Scholar) for how these differ.
  • Google Scholar — the broadest coverage of any source, including grey literature, but with less consistent metadata quality and no straightforward bulk-export mechanism, which limits its use for large-scale, reproducible bibliometric work even though it is a common starting point for smaller literature searches.

Because no two databases index an identical set of journals, conferences, and citation links, the choice of source is itself a methodological decision that should be documented and, where feasible, justified against the research question — a comparison of Web of Science and Scopus results for the same search, for instance, will not produce identical corpora or identical citation counts.

3. Search, extract, and clean the data

A search string is developed and tested against the chosen database(s), then the matching records (with associated metadata: authors, affiliations, journal, publication year, abstract, keywords, references, and citation counts) are exported in bulk, typically as a set of hundreds to many thousands of records. Cleaning is a substantial and often underestimated part of this step: deduplicating records that appear differently across sources, standardizing author names and institutional affiliations (the same author or institution can appear under multiple name variants), and merging near-duplicate keyword or term variants (singular/plural forms, synonyms, abbreviations alongside their spelled-out equivalents). Skipping this step is a common reason a resulting analysis looks noisier or more fragmented than the underlying literature actually is.

4. Select and apply the analytical technique(s)

With a clean corpus in hand, one or more bibliometric techniques (see the taxonomy below) are applied depending on the research question — performance analysis for output and impact, or a science-mapping technique such as co-citation, bibliographic coupling, co-authorship, or co-word analysis for structural questions. Most published bibliometric-analysis papers combine at least a performance-analysis component with one science-mapping technique rather than relying on a single method alone.

5. Visualize and interpret

Network-based techniques are typically rendered as maps — clusters, node size, and layout communicating structure that a table of numbers cannot. VOSviewer and CiteSpace are the two software packages behind most published bibliometric-mapping visualizations; see CASRAI’s Citation Network Analysis guide for how those visualizations are built and interpreted in detail. Interpretation is where the analysis is connected back to the original research question: naming and explaining what each cluster represents, identifying which works or authors are structurally central, and describing how the field has moved over time — a step that requires domain knowledge of the literature, not just software output.

6. Report methodology and limitations

Because two analysts can build meaningfully different results from the same underlying corpus depending on database choice, search string, date range, and analytical thresholds, published bibliometric analyses are expected to report these parameters explicitly so the work is reproducible and its limitations are legible to the reader.

A Taxonomy of Bibliometric Techniques

Bibliometric techniques are commonly grouped into two broad categories: performance analysis, which measures productivity and impact, and science mapping, which reveals the structural and relational aspects of a field. This section is a landscape overview, not a deep dive — several of these techniques have their own dedicated CASRAI coverage linked below.

Performance analysis

  • Citation counts — the number of times a work, author, or journal has been cited; the most basic bibliometric indicator, and the input most other metrics are built from. See CASRAI’s Citation Indexing guide for how citation counts are captured and indexed in the first place.
  • The h-index and related author-level indicators — a single number intended to balance a researcher’s productivity against their impact. See CASRAI’s step-by-step guide to calculating the h-index.
  • Journal- and field-normalized indicators such as the Field-Weighted Citation Impact (FWCI) and the Relative Citation Ratio (RCR), which adjust raw citation counts for the fact that citation norms differ sharply by field and publication age.

Science mapping (relational/network techniques)

  • Citation analysis and citation network analysis — mapping direct citing/cited-by relationships as a network to trace influence and lineage. See CASRAI’s dedicated Citation Network Analysis guide for full coverage of this technique.
  • Co-citation analysis — two works are linked when a third, later work cites both together, revealing which works the field treats as intellectually related. Introduced by Henry Small in 1973; covered in depth in the Citation Network Analysis guide linked above.
  • Bibliographic coupling — two works are linked when they cite the same earlier source in their own reference lists, useful for finding papers addressing similar problems at roughly the same point in time. Introduced by M. M. Kessler in 1963; also covered in the Citation Network Analysis guide.
  • Co-authorship analysis — mapping collaboration structure between researchers, institutions, or countries based on shared authorship, used to reveal who works with whom and which institutions sit at the center of a field’s collaboration network.
  • Co-word (keyword co-occurrence) analysis — mapping which terms or keywords appear together across a corpus’s titles, abstracts, or author-supplied keyword lists, used to surface a field’s dominant themes and how its vocabulary has shifted, independent of citation data.

In practice, most published bibliometric-analysis papers combine several of these techniques rather than applying just one — for example, performance analysis (citation counts, h-index) to establish influential works and authors, alongside co-citation or co-word mapping to establish the field’s structure.

Software and Tools

Once a corpus is exported and cleaned, dedicated software builds and visualizes the maps described above. VOSviewer, developed at Leiden University’s Centre for Science and Technology Studies (CWTS), is free and the most widely used tool in published bibliometric-mapping work, covering co-citation, bibliographic coupling, co-authorship, and keyword co-occurrence in a single point-and-click interface; CASRAI’s VOSviewer guide covers its interface, workflow, and practical usage in full. CiteSpace, from Drexel University, is commonly used alongside VOSviewer with a specific focus on detecting trends and citation bursts over time. The R package Bibliometrix (with its web-app interface, biblioshiny) offers a broader range of statistical bibliometric indicators for those comfortable with a scripting workflow. For readers assembling or exploring a smaller literature set rather than running a full field-level bibliometric study, CASRAI’s guides to Litmaps and lighter citation-mapping tools cover that adjacent, more individual-researcher-focused use case.

Limitations and Responsible Use

  • Database coverage differs. No source indexes an identical set of journals, conferences, and citation links, so a bibliometric analysis run against Web of Science, Scopus, or an open source like OpenAlex will not produce an identical corpus or identical results — this is a methodological choice to document, not a neutral default.
  • Citation counts are not quality judgments. A citation reflects that a work was referenced, not that it was referenced favorably or that its findings were sound; citation-based indicators should be read as measures of influence or attention, not of research quality on their own.
  • Recency bias. Very recent work has had little time to accumulate citations, so citation-based techniques can under-represent genuinely important new research; this is one reason performance metrics are usually paired with a time-aware technique such as co-word analysis or citation-burst detection.
  • Metrics can be gamed. CASRAI’s citation cartel and coercive citation entries describe ways citation data can be manipulated, which is part of why the San Francisco Declaration on Research Assessment (DORA) and the Coalition for Advancing Research Assessment (CoARA) agreement both caution against relying on any single citation-based metric as a standalone measure of quality, whether it comes from performance analysis or from a network’s structural position.

Frequently Asked Questions

What is the difference between bibliometric analysis and a systematic review?

A systematic review follows a documented protocol to identify, screen, and synthesize the substantive findings of individual studies addressing a specific question. Bibliometric analysis instead quantifies patterns across a body of literature — output, citation, and relational structure — and is increasingly used as a complement that maps a field before or alongside a systematic or narrative review, not as a substitute for reading and evaluating individual studies.

What is the difference between bibliometric analysis and citation network analysis?

Citation network analysis is one specific family of bibliometric techniques — mapping citation relationships (co-citation, bibliographic coupling, direct citation) as a network graph. Bibliometric analysis is the broader method that includes citation network analysis alongside performance metrics (citation counts, h-index), co-authorship analysis, and co-word analysis. See CASRAI’s Citation Network Analysis guide for that technique in depth.

Which database should I use for a bibliometric analysis: Web of Science, Scopus, or Dimensions?

There is no universal answer — it depends on the field and question. Web of Science has the deepest historical coverage and is a traditional default in the sciences; Scopus generally has broader journal coverage, especially outside North America and Western Europe; Dimensions adds links to grants, patents, and policy documents. Many rigorous analyses use more than one source, or an open alternative such as OpenAlex, and report the choice explicitly as part of the method.

Do I need programming skills to conduct a bibliometric analysis?

No. Point-and-click tools such as VOSviewer handle the core workflow — import, map creation, visualization — without any scripting. R-based tools such as Bibliometrix offer more statistical depth for those comfortable with a scripting workflow, but are not required to conduct a valid bibliometric analysis.

How large does a corpus need to be for a bibliometric analysis?

There is no fixed threshold, but bibliometric techniques rely on patterns emerging across a body of literature, so they generally need a corpus of at least several dozen to a few hundred records to produce interpretable structure; published field-level bibliometric-review papers frequently analyze corpora in the thousands. A handful of papers is better suited to close reading than to bibliometric mapping.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →