A position paper posted to arXiv in July 2026 does something the growing literature on AI-fabricated references hasn’t done yet: it puts the detection tools themselves under the microscope, rather than just the citations. "Detecting Hallucinated and Suspicious Citations: What Current Tools Can and Cannot Do", by Fidan Badalova and Philipp Mayr (GESIS – Leibniz Institute for the Social Sciences), was submitted 17 July 2026 and revised 28 July 2026. It runs five publicly available hallucinated-citation checkers — CheckIfExist, HalluCiteChecker, Hallucinator, HalRef (Hallucinated Reference Finder) and RefChecker — against documents known to contain fabricated references, and reports where each one succeeds and where it breaks down.
The paper is not a buyer’s guide and does not name a winner. Its core finding is closer to a warning: every tool it tested is limited by some combination of reference-extraction errors, incomplete metadata, narrow database coverage, or inconsistent verification behavior — and none of the five is reliable enough to use unsupervised as a publication gate.
Why this matters now
Editors, research-integrity officers and librarians have been shopping for exactly this kind of tooling for the past year, largely in response to a steady run of high-profile fabricated-citation incidents — see CASRAI’s earlier coverage of the 1-in-277 PubMed/Lancet fabrication-rate audit and of fake-author citation cartels. Until this paper, there was no vendor-neutral comparison of the tools that have sprung up to answer that demand. Badalova and Mayr’s evaluation is the first attempt to test them side by side on a shared set of documents rather than taking each tool’s own reported accuracy at face value.
The five tools, compared
| Tool | What it does | Where it was strong | Where it broke down |
|---|---|---|---|
| RefChecker | Extracts and checks references against bibliographic databases | Flagged more problematic references than most other tools | High false-positive rate — over-flagged legitimate references |
| CheckIfExist | Verifies whether a cited reference actually exists in the literature | Also caught a high proportion of genuinely problematic references | Also produced many false positives, similar to RefChecker |
| HalluCiteChecker | Screens citations for hallucination indicators | Fewer false positives than RefChecker or CheckIfExist | Missed more actually-problematic references (higher false-negative rate) |
| Hallucinator | General hallucinated-reference detector | The most balanced precision/recall trade-off of the five | Not a clear leader on any single measure — a middle performer throughout |
| HalRef (Hallucinated Reference Finder) | Cross-checks reference metadata fields against source records | Produced the most detailed, granular metadata-mismatch signals | Also generated many false positives when metadata records were incomplete |
Read together, the pattern is a familiar precision/recall trade-off: the tools that catch the most real problems (RefChecker, CheckIfExist) also flag the most legitimate references incorrectly, while the tool that flags fewest false alarms (HalluCiteChecker) also lets more genuine problems through. None of the five resolves that trade-off on its own.
Four recurring failure modes
Across all five tools, the paper identifies the same handful of underlying weaknesses recurring in different combinations:
- Reference-extraction errors. Before a tool can verify a citation, it first has to correctly parse the reference list out of the source document. Formatting inconsistencies, unusual citation styles and PDF-extraction artifacts all introduce errors at this first step, and every downstream check inherits them.
- Incomplete metadata. Verification depends on matching a cited reference’s title, authors, venue and year against an external record. When the source document supplies partial metadata — or the underlying bibliographic database’s record is itself incomplete — tools cannot always tell a genuinely fabricated reference from a real one with sparse metadata.
- Limited database coverage. No single bibliographic database indexes every legitimate publication, especially outside English-language, large-publisher, or recent literature. A reference a tool can’t find is not necessarily hallucinated — it may simply sit outside that tool’s search index.
- Inconsistent verification results. The paper reports that the same reference can be scored differently across tools, and in some cases across repeated runs of the same tool — a reliability problem on top of the accuracy problem.
What the paper recommends instead
Badalova and Mayr frame hallucinated and suspicious citations as "a real and growing problem for scientific communication" and argue that no single current tool is sufficient as a standalone gate. Their recommendation is for more transparent, multi-source detection systems — combining several verification signals and disclosing their own uncertainty — rather than treating any one tool’s flag (or clean bill of health) as a final answer. For editors and integrity offices evaluating this class of tooling today, that reads as a case for using these checkers as an early-warning triage layer that narrows down which references a human then verifies, not as an automated pass/fail check that runs unsupervised.
Related CASRAI coverage
- Fabricated citations: what a 1-in-277 PubMed/Lancet audit found
- Fake-author citation cartels
- Dictionary: Hallucination
- Dictionary: Fake Citation
Source
Fidan Badalova and Philipp Mayr, "Detecting Hallucinated and Suspicious Citations: What Current Tools Can and Cannot Do," arXiv:2607.22693, submitted 17 July 2026, revised 28 July 2026. https://arxiv.org/abs/2607.22693







