Turnitin’s AI writing indicator is a statistical estimate of how much a submission’s prose resembles large-language-model output — not a determination that AI was used, and not proof of anything on its own. This guide covers how the detector actually works, what Turnitin’s own accuracy figures and the independent evidence say, what a specific score does and does not mean, and what a defensible process looks like on both sides — for a student or researcher who has been flagged, and for whoever has to decide what the flag is worth.
Does Turnitin check for AI, and who sees the result?
Turnitin’s AI writing detection is a feature bundled into its institutional Similarity and Feedback Studio products — it is not a public or self-serve tool, and there is no way to run it on a document outside an institution that has licensed it. Whether any given submission is actually scored for AI writing depends on whether the instructor’s or institution’s Turnitin account has the feature enabled; it is a separate signal from Turnitin’s long-standing text-matching (similarity/originality) check, and the two can produce completely different results on the same document. A paper can show 0% similarity and still receive a high AI-writing score, because similarity checking looks for matches against a reference database of prior text, while AI-writing detection looks for statistical patterns associated with machine generation. See CASRAI’s Anti-Plagiarism Software guide for how the similarity side works.
How Turnitin’s AI-writing detector actually works
Turnitin’s AI indicator, like most commercial AI-writing detectors, scores text using signals such as perplexity (how predictable each word choice is, given the words before it) and burstiness (how much sentence length and structure vary across a passage). Large language models tend to produce text that is more uniformly predictable and more evenly structured than typical human writing, so the model flags passages that score low on both measures as more likely to be machine-generated. See CASRAI’s Detection tool (AI-generated) entry for the underlying concept.
Critically, this is a statistical estimate, not a fingerprint match. No AI detector, including Turnitin’s, can definitively prove that a specific passage was or was not written by a person, because the signal it relies on — predictability and uniformity — is a matter of degree that also shows up in some kinds of ordinary human writing, discussed below.
Can Turnitin detect ChatGPT specifically?
No — Turnitin’s detector does not identify which tool, if any, produced a piece of text. It has no access to a user’s ChatGPT (or Gemini, Claude, or any other assistant’s) prompt history or session logs, and it does not distinguish between outputs of different language models. What it detects is a general statistical pattern — low word-choice unpredictability and low sentence-length variation — that is characteristic of large-language-model output broadly, whichever model produced it. A result that reads as “Turnitin caught ChatGPT” is, more precisely, “Turnitin’s model judged this passage statistically similar to typical LLM output” — the tool cannot and does not name a specific product, and it has no record of what tool, if any, a writer actually used.
What a Turnitin AI score actually means
If you’re looking at a specific number on a report, there are two things worth knowing: what the percentage is counting, and why some reports show no number at all.
| What you see on the report | What it means |
|---|---|
| Asterisk (*%), no number shown | Turnitin detected some indicators of AI-generated writing, but the estimate fell below its 20% reporting threshold. Turnitin’s own guidance states it does not display a specific percentage or highlight text in this range because its own testing found a higher rate of false positives at low scores. |
| A number from 20% to 100% | The percentage of qualifying text — continuous prose sentences within paragraphs — that Turnitin’s model estimates was generated by a large language model. Headers, quotations, and the reference list are excluded, so the figure is a share of the prose Turnitin actually scores, not a share of the whole document. |
Two things follow from this. First, a report showing *% is not a clean 0% result — the tool found some signal but is deliberately withholding a number because, at that low a signal strength, Turnitin’s own testing says the figure isn’t reliable enough to publish. Second, because the percentage covers only qualifying prose, a document that is largely tables, quotations, code, or a long reference list can show a high percentage on a small base of scored text — worth knowing before treating any single number as a summary of the whole submission.
Is Turnitin’s AI detector accurate?
Two separate bodies of evidence answer this differently, and both matter.
Turnitin’s own published claim: Turnitin states that its AI writing detection achieves roughly 98% accuracy with a false-positive rate under 1%, for documents where more than 20% of the text is AI-generated. That is a vendor-published figure from Turnitin’s own internal testing methodology and conditions — it is not an independent, peer-reviewed measurement, and it explicitly does not describe performance on documents with lower proportions of flagged text, which is exactly the range (the sub-20% asterisk band, discussed above) where Turnitin itself says reliability is weaker.
Independent evidence: The most widely cited independent evaluation is Liang, Yuksekgonul, Mao, Wu, and Zou (Stanford University), published in the journal Patterns in 2023. The study tested seven GPT detectors against 91 TOEFL essays written by non-native English speakers and found an average false-positive rate of 61.3% across those detectors — more than 91% of the essays were flagged by at least one detector — against a near-zero false-positive rate on a control set of essays by native-English-speaking U.S. eighth-graders. The same study found that simple prompt-based rewriting could push genuinely AI-generated text under detection thresholds, meaning the tools it tested were simultaneously over-flagging real human writing from a specific population and under-flagging actual AI text. Later peer-reviewed evaluations in venues including the International Journal for Educational Integrity (Springer) and the journal Information (MDPI) have continued to find inconsistent results across commercial detectors, including Turnitin, on the same text.
Because vendors update their detection models on their own schedule, treat any specific percentage — including the ones above — as a snapshot rather than a permanent figure. For a full evidentiary breakdown written for people who have to weigh detector output as evidence, see CASRAI’s AI Detection Accuracy in Higher Education guide.
Why false positives cluster in specific kinds of writing
The features that make text read as “AI-like” to a detector are, unfortunately, also features that trained academic writing often has. This is why false positives cluster heavily in scholarly and student work rather than spreading evenly:
- Non-native English writing. Second-language academic writers often rely on a narrower, more standard vocabulary and more templated sentence construction, which lowers the same burstiness/perplexity signals detectors use — this is the population the Liang et al. study found was disproportionately flagged.
- Formulaic academic structure. Discipline-standard conventions — topic sentences, transitional phrases, consistent paragraph structure, restrained sentence variation — are exactly the low-burstiness, high-predictability pattern detectors are tuned to catch.
- Heavy editing or professional copyediting. Several rounds of tightening or restructuring for clarity can remove the sentence-length variation and idiosyncratic phrasing that read as “human” to a detector, even when every word was written or approved by the named author.
- Writing to a strict template or style guide. Lab reports, structured abstracts, and methods sections written to a journal’s required format are disproportionately likely to trigger a high score simply because the genre constrains variation.
Institutions rethinking Turnitin’s AI detector
Reliability concerns have led at least some institutions to change how, or whether, they use Turnitin’s AI-writing feature. Curtin University announced it would disable Turnitin’s AI-writing detection from January 1, 2026, while keeping the platform’s text-matching/originality checking in place, framing the decision around trust and clarity in assessment. Claims that circulate online naming a large specific number of universities that have disabled AI detection (commonly cited figures like “50+”) trace back to AI-detector marketing and affiliate content rather than to a verifiable, institution-by-institution tally, and CASRAI is not repeating that figure here. What is verifiable is that reliability concerns are real enough that at least one major research university has acted on them publicly, and that this is an active, evolving policy question rather than a settled one. CASRAI’s comparison of Turnitin’s built-in AI detection against standalone detectors covers how the access model and reported accuracy evidence differ between institutionally bundled and directly purchased tools.
Falsely accused of using AI: what to do
- Don’t panic-rewrite the paper. A flag is a prompt for a conversation, not a verdict. Rewriting everything to “sound less formulaic” can also make a genuinely human paper read as evasive if the matter is later reviewed.
- Check the specific policy that applies to you. Institutional and journal AI policies vary widely in whether, and how, they use AI-writing detection, and what process follows a flag. CASRAI’s guide to how universities are updating academic-integrity policy for AI writing tools covers the current policy landscape; your institution’s or journal’s own written policy is the authoritative source for your situation.
- Gather your process documentation. Draft history in your word processor or reference manager (version history, track-changes logs, timestamped file saves), outlines, notes, and correspondence with advisors or co-authors are the strongest evidence of authentic authorship, because they show a writing process a detector cannot see. Keep this before you need it, not just after a flag.
- Ask what, specifically, was flagged. A whole-document score is less useful than knowing which passages were flagged and why — Turnitin’s report highlights specific sentences rather than issuing a single verdict for the entire document.
- Use the appeal or review process if one exists. Most institutions with a formal academic-integrity process allow a student or author to respond to a flag before any finding is made — present drafts, explain your writing process, and, if relevant, note documented non-native-English-writer status or disability accommodations that may affect writing style.
- Disclose, don’t hide, any legitimate AI assistance. If you used a generative AI tool for a permitted purpose (grammar polishing, translation assistance, brainstorming), most current policies distinguish clearly between disclosed, policy-compliant use and undisclosed use presented as entirely unaided writing. See CASRAI’s Generative-AI disclosure statement entry for what a compliant disclosure typically covers.
What a defensible institutional process requires
For the person deciding what a flagged score is worth — an instructor, an academic-integrity officer, or an editor — the evidence above translates into a small number of concrete practices reflected in a growing body of institutional policy:
- Treat a flag as the start of a review, not its conclusion. A detector score establishes reasonable cause to look further; on its own, it does not establish that misconduct occurred.
- Require corroborating evidence before any finding. Draft history, version-control or track-changes logs, outlines, and correspondence with an advisor are process evidence a detector cannot see and cannot be faked as easily as prose can be edited.
- Give the accused person a right of reply before a finding is made, consistent with standard due-process practice in institutional disciplinary proceedings generally.
- Avoid treating any fixed percentage as a bright-line rule. No detector vendor, and no field-wide standard body, has established a threshold above which a document is definitively AI-generated; a locally adopted cutoff is a policy choice, not a scientific finding, and should be documented and applied consistently.
- Weight false-positive risk by population. Given the documented skew toward non-native English writers and formulaic, heavily edited, or template-constrained writing, a policy that ignores who is disproportionately flagged is not applying the evidence evenly.
CASRAI’s AI Detection Accuracy in Higher Education guide covers this evidentiary framework in full, written specifically for academic-integrity and research-compliance offices.
How this differs from a plagiarism (similarity) flag
It’s worth being precise about which report you’re looking at, because the two are often confused and require different responses. A similarity report (Turnitin Similarity, iThenticate, Crossref Similarity Check) shows what percentage of your text matches existing sources in a reference database — a high score there means investigate potential unattributed copying or citation issues. An AI-writing report shows a statistical estimate of machine-generation likelihood and has nothing to do with matching against other documents. A submission can trigger one, both, or neither, and each requires a different kind of response. See CASRAI’s Anti-Plagiarism Software guide for the similarity-checking side in full.
Frequently asked questions
Does Turnitin check for AI?
Only if the institution’s Turnitin license has the AI writing detection feature enabled — it is a separate, institutionally licensed add-on to Turnitin’s core similarity checking, not a public tool and not automatically active on every account. Whether a specific submission was scored for AI writing depends on that institution’s settings.
How does Turnitin detect AI?
It scores text on statistical signals — primarily perplexity (word-choice predictability) and burstiness (variation in sentence length and structure) — that tend to differ between typical large-language-model output and typical human writing. It does not compare text against a database of known AI outputs the way a similarity checker compares against prior documents.
Can Turnitin detect AI reliably?
Reliably enough to be useful as a starting point for review, not reliably enough to serve as proof on its own. Turnitin’s own published claim is about 98% accuracy with a false-positive rate under 1% for documents where more than 20% of the text is flagged; independent research has found substantially higher false-positive rates in specific populations, particularly non-native English writers. Both figures can be true at once because they describe different test conditions.
Is Turnitin’s AI detector accurate?
It depends heavily on who is being scored and how much of the document is flagged. Turnitin’s own testing reports strong accuracy above its 20% reporting threshold; independent, peer-reviewed testing (notably the Stanford Patterns study) has found detectors in general, including in populations that overlap with tools like Turnitin’s, produce a large false-positive rate for non-native English writers. Neither figure should be treated as the final word without checking your institution’s or journal’s own policy on how it uses a score.
Can Turnitin detect ChatGPT specifically?
No. Turnitin detects statistical patterns associated with large-language-model writing in general — it does not identify which specific AI tool, if any, was used, and it has no access to a user’s prompt or session history with ChatGPT or any other assistant.
What does a Turnitin AI score number actually mean?
It is the percentage of qualifying prose (continuous sentences within paragraphs, excluding headers, quotations, and the reference list) that Turnitin’s model estimates was generated by a large language model. Scores below 20% are not shown as a number at all — they display as an asterisk, because Turnitin’s own testing found a higher false-positive rate at low scores.
What should I do if I’m falsely accused of using AI?
Don’t panic-rewrite your work. Check your institution’s or journal’s specific policy, gather your draft history and version records, ask exactly which passages were flagged, and use any appeal or review process available to you before a finding is made. See the section above for the full process.
Should I disclose AI tool use even if I was only flagged incorrectly?
If you did not use generative AI in a way your institution’s or journal’s policy requires disclosure for, there is nothing to disclose — the flag itself is not evidence of use. If you did use a permitted AI tool (for grammar or translation help, for example) and it wasn’t disclosed, use the flag as a prompt to add a compliant disclosure statement going forward; see the Generative-AI disclosure statement entry for what that typically covers.







