Skip to main content
v2026.11,610 entries · CC-BY 4.0

Editorial · CASRAI · AI and ML research outputs

Five Tools That Check for Hallucinated Citations — and What They Still Miss

A July 2026 arXiv study tests five hallucinated-citation checkers and finds none reliable enough to run unsupervised. What each tool gets right, and wrong.

Published 6 Aug 2026· 4 minute read

Ask about this story

Answers are drawn from this article and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

CASRAI is the reference for research administration — bookmark it for the next question.

A position paper posted to arXiv in July 2026 does something the growing literature on AI-fabricated references hasn’t done yet: it puts the detection tools themselves under the microscope, rather than just the citations. "Detecting Hallucinated and Suspicious Citations: What Current Tools Can and Cannot Do", by Fidan Badalova and Philipp Mayr (GESIS – Leibniz Institute for the Social Sciences), was submitted 17 July 2026 and revised 28 July 2026. It runs five publicly available hallucinated-citation checkers — CheckIfExist, HalluCiteChecker, Hallucinator, HalRef (Hallucinated Reference Finder) and RefChecker — against documents known to contain fabricated references, and reports where each one succeeds and where it breaks down.

The paper is not a buyer’s guide and does not name a winner. Its core finding is closer to a warning: every tool it tested is limited by some combination of reference-extraction errors, incomplete metadata, narrow database coverage, or inconsistent verification behavior — and none of the five is reliable enough to use unsupervised as a publication gate.

Why this matters now

Editors, research-integrity officers and librarians have been shopping for exactly this kind of tooling for the past year, largely in response to a steady run of high-profile fabricated-citation incidents — see CASRAI’s earlier coverage of the 1-in-277 PubMed/Lancet fabrication-rate audit and of fake-author citation cartels. Until this paper, there was no vendor-neutral comparison of the tools that have sprung up to answer that demand. Badalova and Mayr’s evaluation is the first attempt to test them side by side on a shared set of documents rather than taking each tool’s own reported accuracy at face value.

The five tools, compared

Tool What it does Where it was strong Where it broke down
RefChecker Extracts and checks references against bibliographic databases Flagged more problematic references than most other tools High false-positive rate — over-flagged legitimate references
CheckIfExist Verifies whether a cited reference actually exists in the literature Also caught a high proportion of genuinely problematic references Also produced many false positives, similar to RefChecker
HalluCiteChecker Screens citations for hallucination indicators Fewer false positives than RefChecker or CheckIfExist Missed more actually-problematic references (higher false-negative rate)
Hallucinator General hallucinated-reference detector The most balanced precision/recall trade-off of the five Not a clear leader on any single measure — a middle performer throughout
HalRef (Hallucinated Reference Finder) Cross-checks reference metadata fields against source records Produced the most detailed, granular metadata-mismatch signals Also generated many false positives when metadata records were incomplete

Read together, the pattern is a familiar precision/recall trade-off: the tools that catch the most real problems (RefChecker, CheckIfExist) also flag the most legitimate references incorrectly, while the tool that flags fewest false alarms (HalluCiteChecker) also lets more genuine problems through. None of the five resolves that trade-off on its own.

Four recurring failure modes

Across all five tools, the paper identifies the same handful of underlying weaknesses recurring in different combinations:

  • Reference-extraction errors. Before a tool can verify a citation, it first has to correctly parse the reference list out of the source document. Formatting inconsistencies, unusual citation styles and PDF-extraction artifacts all introduce errors at this first step, and every downstream check inherits them.
  • Incomplete metadata. Verification depends on matching a cited reference’s title, authors, venue and year against an external record. When the source document supplies partial metadata — or the underlying bibliographic database’s record is itself incomplete — tools cannot always tell a genuinely fabricated reference from a real one with sparse metadata.
  • Limited database coverage. No single bibliographic database indexes every legitimate publication, especially outside English-language, large-publisher, or recent literature. A reference a tool can’t find is not necessarily hallucinated — it may simply sit outside that tool’s search index.
  • Inconsistent verification results. The paper reports that the same reference can be scored differently across tools, and in some cases across repeated runs of the same tool — a reliability problem on top of the accuracy problem.

What the paper recommends instead

Badalova and Mayr frame hallucinated and suspicious citations as "a real and growing problem for scientific communication" and argue that no single current tool is sufficient as a standalone gate. Their recommendation is for more transparent, multi-source detection systems — combining several verification signals and disclosing their own uncertainty — rather than treating any one tool’s flag (or clean bill of health) as a final answer. For editors and integrity offices evaluating this class of tooling today, that reads as a case for using these checkers as an early-warning triage layer that narrows down which references a human then verifies, not as an automated pass/fail check that runs unsupervised.

Source

Fidan Badalova and Philipp Mayr, "Detecting Hallucinated and Suspicious Citations: What Current Tools Can and Cannot Do," arXiv:2607.22693, submitted 17 July 2026, revised 28 July 2026. https://arxiv.org/abs/2607.22693

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →