A paper is flagged as “AI detected” when detectors such as Turnitin’s AI indicator or GPTZero judge, from statistical signals like low perplexity and burstiness, that its phrasing reads as more machine-typical than human-typical — an estimate that is often wrong for formulaic, precise, or non-native English academic writing. The flag is not proof anything is wrong with your paper.
How AI-detection tools actually work
Most AI-writing detectors, including Turnitin’s, score text using signals such as perplexity (how predictable each word choice is, given the words before it) and burstiness (how much sentence length and structure vary across a passage). Large language models tend to produce text that is more uniformly predictable and more evenly structured than typical human writing, so detectors flag passages that score low on both measures as likely machine-generated. This is a genuinely different technology from the text-matching (originality/similarity) checks covered in CASRAI’s Detection tool (AI-generated) dictionary entry and the Anti-Plagiarism Software guide — a paper can have a 0% similarity score and still receive a high AI-writing score, because the two tools are measuring completely different things.
Critically, these scores are a statistical estimate, not a determination of fact. No AI detector can definitively prove that a specific passage was or was not written by a person, because the underlying signal (predictability and uniformity) is a matter of degree, not a binary fingerprint.
Why formulaic, precise, or non-native academic writing gets flagged
The features that make text look “AI-like” to a detector are, unfortunately, also features that good academic writing is trained to have. This is the core reason false positives cluster so heavily in scholarly and student work rather than being evenly spread across all writing:
- Formulaic academic structure. Discipline-standard conventions — topic sentences, transitional phrases (“furthermore,” “in conclusion,” “it is important to note that”), consistent paragraph structure, and a restrained, low-variance sentence style — are exactly the low-burstiness, high-predictability pattern detectors are tuned to catch. Writers who have been explicitly taught to write “clearly and formally” are, by that same training, writing in a way that resembles the statistical profile of AI output.
- Non-native English writing patterns. Second-language academic writers often rely on a narrower, more standard vocabulary and more templated sentence constructions, which lowers the same burstiness/perplexity signals detectors use. A widely cited 2023 study by Liang, Yuksekgonul, Mao, Wu, and Zou (Stanford University), published in the journal Patterns, tested seven GPT detectors on 91 TOEFL essays written by non-native English speakers and found an average false-positive rate of 61.3% — more than 91% of the essays were flagged by at least one detector — compared with a near-zero false-positive rate on essays by native English-speaking U.S. eighth-graders. The study also found that simple prompt-based rewriting could push AI-generated text under detectors’ thresholds, meaning the same tools were simultaneously flagging real human writing and missing actual AI text.
- Heavy editing, revision, or professional copyediting. A paper that has been through several rounds of tightening, restructuring for clarity, or professional-editing services can lose the sentence-length variation and idiosyncratic phrasing that read as “human” to a detector, even though every word was written or approved by the author.
- Working from a template or a strict style guide. Lab reports, structured abstracts, methods sections written to a journal’s required format, and other highly conventionalized genres are disproportionately likely to trigger a high AI-writing score, simply because the genre itself constrains variation.
What the evidence actually shows about accuracy
This is a fast-moving and contested area, and vendors update their models and stated accuracy claims regularly, so treat any specific percentage — including the ones below — as a snapshot rather than a permanent figure. A few things are well established:
- Independent, peer-reviewed research (the Liang et al. 2023 Patterns study above) found a large, systematic bias against non-native English writers across multiple commercial and research detectors, not an isolated flaw in one product.
- Detector vendors themselves, including Turnitin, publicly caution institutions against treating an AI-writing score as sole or definitive evidence of misconduct, and generally recommend it be used only as one input into a broader academic-integrity conversation, alongside things like drafts, version history, and a conversation with the author — not as an automatic finding of wrongdoing.
- Detection accuracy is reported (by both vendors and independent researchers) to degrade further once AI-generated text has been paraphrased, lightly edited, or run through a “humanizing” tool, and separately, some institutions have moved away from relying on AI-writing scores at all, citing reliability concerns, in favor of process-based or trust-based integrity approaches.
The practical takeaway: an AI-writing percentage on a report is a probabilistic flag worth investigating, not proof of anything on its own — for you as an author, or for whoever is reviewing the report.
What to do if your paper is flagged
- Don’t assume the worst, and don’t panic-rewrite the paper. A flag is a prompt for a conversation, not a verdict. Rewriting everything to “sound less formulaic” can also make a genuinely human paper read as evasive if the matter is later reviewed.
- Check the actual policy that applies to you. Institutional academic-integrity policies and journal/publisher AI policies vary widely in whether, and how, they use AI-writing detection at all, and what process follows a flag. CASRAI’s guide to journal and publisher policies on generative AI in manuscripts covers how editorial offices are currently handling AI-related disclosure and review; your own institution or target journal’s written policy is the authoritative source for your specific situation, not general guidance like this page.
- Gather your process documentation. Draft history in your word processor or reference manager (version history, track-changes logs, timestamped file saves), notes, outlines, and correspondence with advisors or co-authors are the strongest evidence of authentic, human authorship, because they show the writing process a detector cannot see. Keep this documentation before you need it, not just after a flag.
- Ask what, specifically, was flagged. A whole-document score is much less useful than knowing which passages were flagged and why — most detector reports (including Turnitin’s) highlight specific sentences or sections rather than issuing a single yes/no verdict for the entire document.
- Use the appeal or review process if one exists. Most institutions with formal academic-integrity processes allow a student or author to respond to a flag before any finding is made — treat this as your opportunity to present drafts, explain your writing process, and, if relevant, note documented non-native-English-writer status or disability accommodations that may affect writing style.
- Disclose, don’t hide, any legitimate AI assistance. If you did use a generative AI tool for a permitted purpose (grammar polishing, translation assistance, brainstorming), most current policies distinguish clearly between disclosed, policy-compliant use and undisclosed use presented as entirely your own unaided writing. See CASRAI’s Generative-AI disclosure statement entry for what a compliant disclosure typically covers, and the AI-generated content and Generative AI dictionary entries for how these terms are defined.
- If you want to understand how the detector itself works, compare tools before treating any one score as definitive. Detectors differ in what they measure and how they report results — CASRAI maintains an independently researched comparison of Originality.ai and GPTZero, part of its broader guide to AI writing and detection tools. (Disclosure: some links on those pages are CASRAI referral links.)
How this differs from a plagiarism (similarity) flag
It’s worth being precise about which report you’re looking at, because the two are often confused and require different responses. A similarity report (Turnitin Similarity, iThenticate, Crossref Similarity Check) shows what percentage of your text matches existing sources in a reference database — a high score there means investigate potential unattributed copying or citation issues. An AI-writing report shows a statistical estimate of machine-generation likelihood and has nothing to do with matching against other documents. A paper can trigger one, both, or neither, and each requires a different kind of response. See CASRAI’s Anti-Plagiarism Software guide for the similarity-checking side of this in full.
Frequently asked questions
Can Turnitin’s AI detector be wrong?
Yes. Independent research has documented systematic false positives, particularly for non-native English writers and formulaic academic writing, and Turnitin itself cautions against treating a score as definitive proof. Detector accuracy also varies by how much a text has been edited or paraphrased.
Does paraphrasing or heavy editing trigger an AI-detected flag?
It can. Heavy revision, professional copyediting, and writing to a strict template or style guide all tend to reduce the sentence-length variation and word-choice unpredictability that detectors use as their main human-versus-AI signal, which can push genuinely human writing toward a higher AI-writing score.
What percentage AI score is considered a problem?
There is no universal threshold — it depends entirely on your institution’s or journal’s written policy, and those policies vary. Some explicitly instruct reviewers not to treat any single percentage as an automatic finding; check the specific policy that applies to your submission rather than assuming a number from another institution applies.
Should I disclose AI tool use even if I was only flagged incorrectly?
If you did not use generative AI in a way your institution’s or journal’s policy requires disclosure for, there is nothing to disclose — the flag itself is not evidence of use. If you did use a permitted AI tool (e.g., for grammar or translation help) and it wasn’t disclosed, use the flag as a prompt to add a compliant disclosure statement going forward; see the Generative-AI disclosure statement entry for what that typically covers.







