Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

How to Detect AI-Generated Text in Academic Writing

A practical, multi-signal guide to spotting AI-generated text in academic work: stylistic tells, citation fabrication, perplexity/burstiness detection, and tool limitations.

Detecting AI-generated text in academic writing is not a single test — it is a combination of close reading, citation checking, and (cautiously) statistical tools, none of which is reliable enough to use alone. Editors, reviewers, supervisors, and integrity offices increasingly need to make a defensible judgment about whether a submission was substantially AI-written, but every method below has real error rates. This guide covers the actual mechanics of detection: what to look for by eye, how algorithmic detectors work and where they fail, and how institutions and journals are expected to use this evidence responsibly.

Before using any of these methods to make an accusation, read CASRAI’s guide on why AI-detection tools produce false positives — the same statistical signals that flag AI writing also flag some kinds of human writing, including work by non-native English speakers and writing in narrow technical registers.

Manual, stylistic signals

Before reaching for a tool, most experienced editors and instructors notice a cluster of stylistic tells. No single one is proof, and all of them occur in ordinary human writing too — but several together, especially combined with the citation problems below, are worth a closer look.

  • Uniform sentence rhythm. Large language models tend to produce sentences of similar length and structure across a passage, without the natural variation (“burstiness”) of human prose — see the statistical detail below.
  • Generic hedging and false balance. Repeated “on the one hand… on the other hand” framing, unearned qualifiers (“it is important to note that,” “in conclusion, it is clear that”), and conclusions that restate the introduction almost verbatim.
  • Confident but shallow claims. Statements that are fluent and grammatically correct but vague where a genuine expert would be specific — missing dataset names, sample sizes, instrument settings, or field-specific caveats a practitioner would naturally include.
  • Inconsistent voice across sections. A methods section that reads fluently but generically, next to a results section with real specificity, can indicate mixed authorship — partly drafted, partly AI-generated or AI-polished.
  • Overuse of certain transitional and evaluative words. Words like “delve,” “boast,” “underscore,” “tapestry,” and “furthermore” appear disproportionately often in outputs from several widely used models, though this signal has weakened as later model versions and prompting have diversified vocabulary — treat it as a mild prompt to look closer, not evidence on its own.

Check the citations

Citation checking is one of the highest-value, lowest-tech detection steps, because it catches a failure mode that is specific to generative AI rather than shared with ordinary human writing: fabricated references. Large language models can generate citations that look completely plausible — real author names, a real journal, a plausible year — for papers that do not exist, or that misattribute real findings to the wrong source.

This is not a hypothetical risk. Research by Walters and Wilder (2023, Scientific Reports) found that a substantial share of citations generated by ChatGPT-3.5 and ChatGPT-4 across a large sample of prompts were fabricated. Springer Nature retracted a 2025 machine-learning textbook after Retraction Watch found that roughly two-thirds of a sampled set of its citations either did not exist or contained substantial errors. Retraction Watch’s 2026 analysis reported that detections of fabricated references in PubMed-indexed literature rose roughly twelve-fold over two years — a trend consistent with wider generative-AI use in manuscript drafting.

Practical check: pull a sample of citations — not just the first few — and verify each one actually exists (DOI resolves, journal/volume/page match) and actually supports the claim it is attached to. A reference list with several unfindable entries, or entries that exist but say something different from what the text claims, is a strong signal regardless of what any detector reports. See CASRAI’s AI-generated content and paper mills and tortured phrases guide for related integrity red flags that often co-occur with fabricated references.

Statistical detection: perplexity and burstiness

Most automated AI-writing detectors — including Turnitin’s AI writing indicator and standalone tools like GPTZero — score text using two related statistical signals:

  • Perplexity measures how predictable each word choice is given the words before it. Language models are trained to produce the statistically likely next word, so AI-generated text tends to score lower on perplexity (more predictable) than typical human writing.
  • Burstiness measures how much sentence length and structure vary across a passage. Human writing tends to alternate between short and long sentences, simple and complex structures; model output tends to be more uniform.

Detectors combine these signals (often alongside other features) into a probability score, not a binary yes/no answer. That distinction matters: a “78% AI-generated” result is a statistical estimate with a real error rate in both directions, not a forensic finding. See CASRAI’s Detection tool (AI-generated) dictionary entry and the false-positives guide for how these scores get misread in practice, including why heavily-edited human writing, translated text, and formulaic technical writing (methods sections, legal boilerplate) can score as “likely AI” without being AI-generated at all.

Automated detection tools

Two broad categories of tool are in use in academic settings:

  • Institutional plagiarism-suite add-ons, such as Turnitin’s AI writing indicator, which most universities already license as part of a text-matching subscription and which integrates directly with the LMS submission workflow. These are covered in CASRAI’s Anti-Plagiarism Software guide, which distinguishes text-matching (originality/similarity checking against existing sources) from AI-writing detection — a genuinely different technology that estimates likelihood of machine generation rather than matching to a database.
  • Standalone AI detectors, such as GPTZero and Originality.ai, marketed directly to educators, editorial offices, and integrity teams, often with higher per-document throughput or API access for bulk screening.

Independent published accuracy assessments of these tools have found meaningful false-positive and false-negative rates, and performance that degrades against text that has been paraphrased, machine-translated, or lightly “humanized” after AI drafting — none of the major detectors claims courtroom-grade certainty, and most vendors themselves recommend the score be treated as one input to a human decision, not the decision itself.

For a side-by-side look at two widely used standalone detectors, see CASRAI’s independently researched comparison, Originality.ai vs GPTZero, and the individual reviews of Originality.ai and GPTZero. (Disclosure: some links on those pages are CASRAI referral links; the comparison itself explicitly recommends testing any detector against your own documents before relying on it, rather than treating either as a sole basis for an integrity decision.)

How journals and institutions are expected to use this evidence

Publication-ethics bodies have been explicit that AI tools cannot be listed as authors and that any generative AI used in drafting text, analyzing data, or producing figures must be disclosed. The Committee on Publication Ethics (COPE) stated this directly in its February 2023 position statement Authorship and AI tools: AI tools cannot take responsibility for a work, cannot hold copyright, and cannot manage conflicts of interest, so they cannot qualify as authors — but their use, and how it was used, should be disclosed, typically in the methods or acknowledgments section. The International Committee of Medical Journal Editors (ICMJE) has taken a parallel position: chatbots and similar tools do not meet ICMJE’s authorship criteria and their use in manuscript preparation should be disclosed.

Neither body endorses using an AI-detection score as standalone proof of misconduct. In practice, journals and universities that act on detection results tend to treat a flagged score as a trigger for a closer human review — checking citations, comparing against a known writing sample, and asking the author directly — rather than as a finding in itself. That distinction is what separates a defensible integrity process from one that is vulnerable to challenge, given the documented false-positive rate of every current detector.

A practical, multi-signal workflow

  1. Read for stylistic tells first — uniform rhythm, generic hedging, shallow specificity, voice inconsistency across sections.
  2. Spot-check citations — verify a real sample resolves and actually supports the claims attached to it, not just that a reference list exists.
  3. Run a detector if your institution licenses one, and treat the output as a probability, not a verdict — note which detector, what score, and what threshold your institution’s policy uses.
  4. Compare against a known writing sample where one exists (a prior draft, an in-class writing sample, earlier published work) — genuine stylistic drift is more informative than any single detector score.
  5. Ask before you accuse. Every major publication-ethics and academic-integrity body that has published guidance on this recommends a conversation with the author before a formal finding, precisely because of the false-positive rate documented in independent testing.

Frequently asked questions

Can AI detectors be fooled?

Yes. Paraphrasing tools, “humanizer” services, and simple manual editing (varying sentence length, replacing high-frequency AI vocabulary) can reduce a detector’s confidence score substantially. This is a known limitation of the underlying statistical approach — it measures patterns typical of unedited model output, and those patterns are exactly what light editing removes. It is one reason no single detector score should be treated as conclusive.

Do AI detectors work on languages other than English?

Most widely used detectors were trained primarily on English text and have published or independently tested accuracy that is lower, and false-positive rates that are higher, on non-native-English and translated writing. This is one of the most consistently documented failure modes across independent testing and a major reason detector output alone should never be the basis for an accusation against a specific student or author.

Is using AI to help write an academic paper automatically misconduct?

No — using generative AI for tasks like language polishing, literature search assistance, or drafting help is not automatically prohibited by most journals or institutions, but it is very often subject to a disclosure requirement (see the COPE and ICMJE positions above). The undisclosed use of substantial AI-generated text presented as one’s own original writing, or AI-fabricated content such as invented citations or data, is the actual integrity concern — not AI assistance itself. See CASRAI’s guide on whether ChatGPT-generated text counts as plagiarism for how this distinction is typically drawn.

What should I do if my own writing was flagged as AI-generated?

See CASRAI’s dedicated guide, Why Does My Paper Say “AI Detected”?, which covers why false positives happen, what the research shows about their frequency, and what to do next.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →