Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Citation Management with AI Writing Tools: Features and Accuracy Risks

How AI writing tools like Jenni.ai and Paperpal handle citation suggestion and formatting, how they relate to Zotero, EndNote, and Mendeley, and the fabricated-citation risk researchers need to verify.

Citation management is no longer just the job of a standalone reference manager. Over the last two years, AI writing assistants built for researchers — tools such as Jenni.ai and Paperpal — have started to fold citation-related features directly into the drafting process itself: suggesting sources as a sentence is written, flagging formatting inconsistencies against a target style, and in some cases syncing straight into an existing Zotero or Mendeley library. This guide covers what these AI-powered citation features actually do, how they differ from (and connect to) the reference managers researchers already use, and the specific accuracy risk — AI-fabricated citations — that makes independent verification non-negotiable before anything reaches a submitted manuscript.

What counts as “AI-powered citation management”

Three distinct capabilities tend to get lumped under this label, and it is worth separating them because they carry different levels of risk:

  • Citation suggestion and relevance ranking — the tool searches academic databases or a user’s uploaded library as the manuscript is drafted and proposes sources that plausibly support a given sentence or claim.
  • Automated formatting-error detection — the tool checks in-text citations and a reference list against a target citation style (APA, MLA, Chicago, Vancouver, and others) and flags inconsistencies such as a missing DOI, an incorrect author-count truncation, or a mismatched in-text/reference-list entry.
  • Citation generation by a general-purpose language model — a chatbot-style AI is asked directly to produce a reference or bibliography, drawing on its trained knowledge rather than a live, verifiable database lookup. This is the category with the highest documented error rate, discussed below.

The first two are increasingly built into dedicated academic writing tools with access to real bibliographic databases; the third is what happens when a researcher treats a general-purpose chatbot as if it were a citation database, which it is not.

Citation suggestion and relevance ranking

Tools built specifically for academic writing — Jenni.ai is a widely used example — search live academic sources or literature APIs as a user drafts, then rank candidate citations by apparent relevance to the surrounding sentence, similar in spirit to how a search engine ranks results. Because the underlying lookup queries a real database rather than generating text from the model’s trained parameters, this class of feature is meaningfully less prone to outright fabrication than asking a chatbot to “write me a citation” — but relevance ranking is still a judgment call the software is making, not the researcher, and a source can be topically related without actually supporting the specific claim it gets attached to. The output is a starting point for the author to check, not a citation the author can insert unread.

Automated formatting-error detection

The second common feature is a style-compliance check: scanning a manuscript’s in-text citations and reference list against the conventions of a specified citation style and surfacing likely errors — inconsistent author formatting, missing publication years, a reference cited in text but absent from the bibliography (or vice versa), or a style mismatch against a journal’s stated requirements. This is closer in spirit to a spell-checker than to a research tool: it is checking mechanical consistency against a known rule set, not verifying that a citation is real or that it actually supports the claim near it. For a full breakdown of the individual citation-style conventions this kind of check is validating against, see CASRAI’s guide to citation styles and referencing formats.

Where Jenni.ai and Paperpal fit

Jenni.ai and Paperpal are two of the more widely used AI academic-writing assistants that have added citation-related functionality on top of their core drafting/editing tools, and each has taken a different emphasis. Jenni.ai’s citation feature searches academic sources and inserts a formatted citation at the point of writing, with support for a large number of citation styles, and can import and export via RIS so a working library can move to or from Zotero or Mendeley rather than being locked inside the tool. Paperpal is built more around manuscript polish ahead of journal submission — grammar, style, and a battery of pre-submission technical checks — with citation-style support layered in as part of that broader submission-readiness role. Neither tool is a full replacement for a dedicated reference manager’s library-management, group-sharing, and long-term archival features (see CASRAI’s Zotero vs EndNote vs Mendeley comparison for that comparison in depth); they are better understood as citation-aware layers on top of the drafting process, feeding into or pulling from a reference manager rather than replacing one. This page does not carry any referral or affiliate relationship with either product — mentions here are descriptive, not a recommendation of one over the other.

The real risk: AI-fabricated citations

The accuracy risk in this space is well documented and specific to one category above: asking a general-purpose large language model to produce a citation directly, from its trained knowledge, rather than through a live database lookup. A 2023 study in Scientific Reports (Walters & Wilder, “Fabrication and errors in the bibliographic citations generated by ChatGPT”) found that a substantial share of references generated this way did not correspond to any real publication, and that even citations pointing to genuinely existing works often contained errors in author names, dates, or other details. Independent follow-up studies since have found fabrication rates that vary considerably by model, topic, and prompt — some newer models perform better than early ChatGPT versions, but no widely used model has been shown to eliminate the problem, and researchers in some subfields have reported fabrication rates as high as one in four or worse. The practical takeaway is not that every AI citation tool is unreliable — tools that search a live database and return a real, retrievable source (the suggestion-and-ranking category above) are a different and generally safer category than a chatbot inventing a reference from memory — but that no AI-suggested or AI-generated citation should reach a manuscript’s reference list without the author independently confirming the source exists, that the title/author/year/DOI are correct, and that it actually says what the citation claims it says.

A verification checklist before submission

  • Open every AI-suggested or AI-generated citation at its source (publisher page, DOI, or database record) rather than trusting the tool’s formatted output.
  • Confirm the DOI resolves and matches the cited title, authors, and year.
  • Read enough of the actual source to confirm it supports the specific claim it is attached to, not just that it is topically related.
  • Run the finished reference list through the target citation style’s actual requirements (or a formatting-check tool built on a real style database), rather than trusting a general AI assistant’s formatting from memory.
  • Keep the verified library in a dedicated reference manager (Zotero, EndNote, or Mendeley) as the system of record, using an AI writing tool’s citation feature as a drafting aid that feeds into it, not as a replacement for it.

Do AI citation features need to be disclosed to a journal?

Most current publisher AI-disclosure policies (ICMJE, COPE, and individual publishers such as Elsevier and Springer Nature) draw their disclosure line around content generation and substantive rewriting, not around using a database-backed citation-formatting or suggestion tool — the distinction is closer to “did the AI produce or alter scientific content” than “did software help format a bibliography.” That said, disclosure norms in this area are moving quickly and vary by publisher, and using a general-purpose chatbot to generate citation text (as opposed to a database-backed lookup) is exactly the kind of AI-generated content most current policies expect to be disclosed. See CASRAI’s dictionary entry on the generative-AI disclosure statement and the related guide on AI proofreading tools, which covers the broader disclosure line between language polishing and content generation in more depth, and always check the specific journal’s current author instructions rather than relying on a general rule.

Frequently asked questions

Can I trust an AI tool’s citation suggestions without checking them?

No. Even tools that search a real academic database rather than generating a citation from a language model’s memory are making a relevance judgment, not a factual guarantee that the source supports the specific claim it is attached to. Every AI-suggested citation should be opened and read before it is kept in a manuscript’s reference list.

Is it safer to use a chatbot or a dedicated academic writing tool for finding citations?

A tool that performs a live lookup against an academic database (the model most AI academic-writing assistants use for their citation-suggestion feature) is generally safer than asking a general-purpose chatbot to produce a citation directly from its trained knowledge, which is the pattern most strongly associated with fabricated references in published research on the topic.

Do Jenni.ai or Paperpal replace Zotero, EndNote, or Mendeley?

Not fully. Both add citation-aware features to the drafting process, and Jenni.ai in particular supports RIS import/export to move references to and from a dedicated reference manager, but none of the AI writing assistants match a dedicated reference manager’s full library management, PDF annotation, and long-term collection features. See CASRAI’s Zotero vs EndNote vs Mendeley comparison for that comparison.

Does using an AI citation-formatting tool need to be disclosed to a journal?

Most current publisher policies do not require disclosure for mechanical formatting or database-backed citation lookup, distinguishing it from AI-generated scientific content. Policies vary and are evolving quickly, so check the specific journal’s current author instructions, and see CASRAI’s dictionary entry on the generative-AI disclosure statement for the general framework publishers are applying.

For the broader landscape of AI tools in manuscript preparation, see CASRAI’s Scholarly Writing pillar page.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →