Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Citation Indexing: What It Is and How Citation Indexes Work

A foundational explainer on citation indexing: how citation indexes extract, match, and count citations, how Web of Science, Scopus, and Google Scholar differ, and how citation data is used (and misused) in research assessment.

Citation indexing is the practice of systematically recording, for every item in a document collection, the other items it cites — and then organizing that data so it can be searched in reverse: starting from a paper, you can find every later paper that cited it. That reverse-lookup capability (forward-citation searching) is what distinguishes a citation index from an ordinary bibliography or subject index, and it is the underlying data layer that most modern research-assessment metrics, discovery tools, and bibliometric analyses are built on top of.

This page covers the general concept — what a citation index actually is, how the major citation indexes are built and differ from one another, and how citation counts and citation networks get used (and misused) in research assessment. For deep dives on specific platforms and metrics, see the related CASRAI pages linked throughout, including Web of Science, Scopus, Journal Impact Factor, and h-index.

What is a citation index?

A citation index is a database that, for each indexed publication (the “source item”), captures its full reference list (the “cited references”) as structured, searchable data rather than as plain text in a bibliography. Because every reference is captured as a discrete, matchable record, the database can be queried in the direction a traditional library catalog never supported: given a specific paper, book, or patent, which later works cited it? That operation — forward-citation search — is the defining feature of a citation index, as opposed to a subject index (which groups items by topic) or a simple bibliography (which only lists what a given work cites, not what later cites it).

The concept was proposed by information scientist Eugene Garfield in 1955, in a paper arguing that citation data could be used to search the literature and to study its structure. Garfield founded the Institute for Scientific Information (ISI) in Philadelphia in 1960, and ISI published the first Science Citation Index (SCI) in 1964 — described by its current owner, Clarivate, as the world’s first citation index. ISI extended the same model to other fields over the following years: the Social Sciences Citation Index (SSCI) launched in 1973 and the Arts & Humanities Citation Index (AHCI) in 1978. Journal Citation Reports (JCR), which introduced the Journal Impact Factor, launched in 1976, built directly on the same underlying citation data.

ISI’s citation indexes changed corporate ownership several times: Thomson Corporation acquired ISI in 1992; Thomson merged with Reuters in 2008 to form Thomson Reuters; in 2016, Thomson Reuters’ scientific and academic research division — including ISI and the Web of Science platform — was spun out and sold, and rebranded as Clarivate (initially Clarivate Analytics). The online platform itself was launched in 2002 as “ISI Web of Knowledge,” continued as “Thomson Reuters Web of Knowledge” after the 2008 merger, and was renamed “Web of Science” effective January 5, 2014 — the name it has carried since.

How citation indexing works

Every citation index performs the same basic set of operations, whatever the underlying technology:

  1. Source-item selection. The index defines a scope of journals, conference proceedings, books, or other document types it will cover, and a process for deciding what gets added or removed. This selection process is the single biggest driver of differences between citation indexes — it determines what counts as “in” the citation network at all.
  2. Reference extraction. For every indexed item, the database parses its full reference list into structured, individually matchable records (author, title, source, year, volume, pages, and — increasingly — a DOI or other persistent identifier) rather than leaving it as unstructured text.
  3. Reference matching (citation linking). Each extracted reference is matched against the index’s own catalog of source items. When a reference in Paper B’s bibliography matches Paper A’s record, the index creates a citation link from B to A. This step is where indexing quality varies most in practice — inconsistent formatting, abbreviated journal names, and metadata errors all cause missed or incorrect matches.
  4. Citation counting and network construction. Once links exist in both directions, the index can report, for any item, how many later items cite it (a “times cited” count) and can trace multi-step citation paths — building a citation network or citation graph in which papers are nodes and citations are directed edges. This same graph structure underlies techniques sometimes discussed under the separate heading of citation network analysis, which applies graph-analytic methods (co-citation clustering, bibliographic coupling, network visualization) to the data a citation index provides — citation indexing supplies the underlying dataset; citation network analysis is one family of things researchers do with it.

The major citation indexes today

Web of Science (Clarivate)

Web of Science is Clarivate’s citation-indexing platform, built around the Web of Science Core Collection — a bundle of six constituent indexes: the Science Citation Index Expanded (SCIE), the Social Sciences Citation Index (SSCI), the Arts & Humanities Citation Index (AHCI), the Emerging Sources Citation Index (ESCI, for journals under evaluation that don’t yet meet full SCIE/SSCI/AHCI selection criteria), the Conference Proceedings Citation Index, and the Book Citation Index. Coverage is editorially curated: journals are evaluated against published selection criteria before inclusion, which is why Web of Science coverage is narrower but more consistently vetted than a general web-crawl-based index. Journal Citation Reports and the Journal Impact Factor are both derived from Web of Science citation data. See CASRAI’s guides on searching Web of Science effectively, the Web of Science Citation Report, and the Web of Science Master Journal List.

Scopus (Elsevier)

Scopus is Elsevier’s competing citation and abstract database, curated by an independent Content Selection and Advisory Board (CSAB) against its own published inclusion criteria, distinct from Web of Science’s. Scopus counts citations only from documents indexed within Scopus itself, and produces its own family of metrics — CiteScore, SJR (SCImago Journal Rank), and SNIP (Source Normalized Impact per Paper) — separate from the Journal Impact Factor. See CiteScore vs. Journal Impact Factor and how to create a Scopus Author ID for more detail, and the direct comparison at Google Scholar vs. Scopus.

Google Scholar

Google Scholar is a free, automated search engine over scholarly literature that also functions as a citation index: it extracts and matches references much as Web of Science and Scopus do, but with a fundamentally different selection process. Rather than editorial curation against published criteria, Google Scholar crawls the web and indexes whatever scholarly-looking documents its algorithms find — journal articles, theses, preprints, conference papers, patents, books, and grey literature such as reports, slide decks, and working papers. This produces substantially broader coverage and, in the large majority of comparative studies, equal-or-higher citation counts than Scopus or Web of Science for the same item, because Google Scholar counts citations from a wider range of document types with no independent gatekeeping. The tradeoff is that Google Scholar has no published inclusion criteria, no official public API, and is comparatively easy to manipulate — a limitation discussed further below. See CASRAI’s guide to creating and optimizing a Google Scholar profile and the comparison at Google Scholar vs. Web of Science.

Open and emerging alternatives

A newer generation of open, algorithmically-built citation databases has grown alongside the three incumbents above, most notably OpenAlex, which indexes scholarly works, authors, and citation links as fully open data (successor in spirit to the discontinued Microsoft Academic Graph), along with Dimensions and Semantic Scholar. CASRAI’s metadata search engines comparison and the three-way Scopus vs. Web of Science vs. OpenAlex comparison cover how these differ in coverage, licensing, and access model from the subscription incumbents.

How citation counts and networks are used in research assessment

Citation-index data underpins most quantitative research-assessment metrics in current use:

  • The Journal Impact Factor (JIF) and Journal Citation Reports rank journals by average citations per article over a defined window, using Web of Science data.
  • The h-index and related author-level metrics summarize an individual researcher’s citation record, computable from any of the major indexes (with different results depending on which index’s coverage and counting rules are used).
  • The Relative Citation Ratio (RCR), developed by NIH, and the Field-Weighted Citation Impact (FWCI), used in Scopus-derived tools, both attempt to normalize raw citation counts for field and publication-age differences, since citation practices vary enormously across disciplines.
  • Citation networks built from indexed reference data are also used more directly, for literature mapping and discovery (finding related work by following citation links forward and backward) rather than strictly for scoring — tools built for this purpose are covered in CASRAI’s guides to Litmaps and ResearchRabbit.

Because these metrics carry real institutional weight — in hiring, promotion, tenure, and funding decisions — two influential initiatives argue for using citation-derived metrics carefully rather than as a substitute for expert judgment: the San Francisco Declaration on Research Assessment (DORA), launched in 2012, calls specifically for not using journal-level metrics such as the Journal Impact Factor as a proxy for the quality of an individual researcher’s or article’s contributions; and the Leiden Manifesto, published in 2015, sets out ten principles for the responsible use of bibliometrics in research evaluation, including that quantitative data should support, not replace, qualitative expert assessment.

Limitations and known issues

A few structural limitations are worth understanding before treating citation-index data as an objective measure of quality or impact:

  • Coverage differs by index. Because Web of Science and Scopus apply editorial selection and Google Scholar does not, the same paper can show meaningfully different citation counts across the three — comparisons across indexes should account for this rather than treating “citation count” as a single, portable number.
  • Field and document-type bias. Citation practices vary widely by discipline (fast-moving experimental fields cite more, and sooner, than mathematics or the humanities) and by document type (a review article typically accumulates more citations than an original research article), which is why normalized metrics like RCR and FWCI exist.
  • Gaming and manipulation. Because citation counts feed directly into metrics with real career and funding consequences, they create an incentive to manipulate them — through practices such as citation cartels (coordinated groups that cite each other disproportionately), coercive citation (editors or reviewers pressuring authors to add citations), excessive self-citation, and outright fake citations. Broader, less curated indexes such as Google Scholar are generally considered more vulnerable to this kind of manipulation than editorially-curated indexes, precisely because there is no independent gatekeeping step before an item is indexed.
  • A citation is not automatically an endorsement. Raw citation counts don’t distinguish between a citation that builds on a work’s findings and one that cites it only to critique, correct, or disagree with it — a limitation some tools (such as Scite’s “Smart Citations”) attempt to address by classifying the nature of each citing statement rather than treating every citation as equivalent.

Citation indexing vs. citation network analysis

These two terms are closely related but not interchangeable. Citation indexing is the underlying infrastructure: the process of extracting, matching, and storing citation links so that “who cited this?” and “what does this cite?” can be answered as a database query. Citation network analysis is a family of analytical techniques applied on top of that data — using graph-analytic methods such as co-citation analysis and bibliographic coupling to map how a body of literature is structured, identify influential works or clusters, or trace the intellectual lineage of a research area. In short: a citation index is the dataset; citation network analysis is one of several things a researcher can do with it. Author-level and journal-level metrics (h-index, Journal Impact Factor) are a different, simpler use of the same underlying data — summary statistics rather than network analysis.

Frequently asked questions

Is Google Scholar a citation index?

Functionally, yes — it extracts and matches references and supports forward-citation search the same way Web of Science and Scopus do. What sets it apart is its selection process: Google Scholar has no published inclusion criteria and relies on automated web crawling rather than editorial curation, which produces broader coverage but less consistent quality control.

What is the difference between a citation index and a database like PubMed?

A bibliographic database like PubMed indexes publications by subject and metadata for discovery and retrieval, but not every bibliographic database captures and links reference lists the way a citation index does. A citation index specifically supports the reverse lookup — finding what later cited a given work — which requires structured, matched reference data as its core feature.

Which citation index should I use for research assessment?

There is no single correct answer; the major frameworks for responsible metric use (DORA, the Leiden Manifesto) generally advise using citation data as one input alongside expert judgment rather than a sole basis for decisions, and being explicit about which index and which counting window a given number comes from, since the same work can show different counts across Web of Science, Scopus, and Google Scholar.

How far back do citation indexes go?

This varies by index and by which constituent index within a platform is being used — Web of Science’s Science Citation Index Expanded, for example, has backfile coverage extending to 1900 within its current electronic edition, while other constituent indexes and other platforms have different start dates and depths. Check the specific index’s own coverage documentation rather than assuming a uniform start date.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →