Academic-integrity and research-compliance offices increasingly encounter AI-text-detection output — a Turnitin AI writing score, a GPTZero result, an Originality.ai report — as evidence in a potential misconduct case. This guide is written for that office: what “accuracy” actually means for these tools, what the published research shows about their real-world error rates, and why a growing body of institutional guidance treats detector output as one input into an investigation rather than a stand-alone finding. It is a companion to CASRAI’s student-facing explainer on AI-detection false positives and its broader guide to how universities are updating AI academic-integrity policy; this page focuses specifically on the evidentiary weight a detector score can reasonably bear.
What “accuracy” means for an AI-text detector
AI-writing detectors do not compare a submission against a reference database the way a similarity checker does (see CASRAI’s Anti-Plagiarism Software guide for that distinction). Instead they estimate, from statistical signals like word predictability (perplexity) and sentence-length variation (burstiness), how likely a passage is to have been produced by a language model. Two error types matter for an integrity office, and they trade off against each other:
- False positive — the tool flags human-written text as AI-generated. This is the error type with direct due-process consequences: it can put a genuine author’s work under investigation.
- False negative — the tool fails to flag text that was, in fact, AI-generated, particularly after paraphrasing or light editing.
A single “accuracy” percentage from a vendor usually blends both error types under specific test conditions (document length, how much of the text is AI-generated, which model produced it). Those conditions rarely match a real submission, which is one reason vendor-reported accuracy and independently measured accuracy diverge as often as they do.
What the published research shows
Vendor claims
Turnitin has publicly stated that its AI writing detection achieves roughly 98% accuracy with a false-positive rate under 1%, for documents where more than 20% of the text is AI-generated. That figure describes Turnitin’s own internal testing methodology and conditions; it is not an independent, peer-reviewed measurement, and it does not describe performance on documents with lower proportions of AI-generated text or the specific case of a human-written document that merely resembles AI output stylistically.
Independent findings on non-native-English writers
The most widely cited independent evaluation is Liang, Yuksekgonul, Mao, Wu, and Zou (Stanford University), published in the journal Patterns in 2023. The study ran seven GPT detectors against 91 TOEFL essays written by non-native English speakers and found an average false-positive rate of 61.3% across those detectors — more than 91% of the essays were flagged by at least one detector — against a near-zero false-positive rate on a control set of essays by native-English-speaking U.S. eighth-graders. The same study found that simple prompt-based rewriting could push genuinely AI-generated text under detection thresholds, meaning the tools were simultaneously over-flagging real human writing from a specific population and under-flagging actual AI text.
OpenAI’s own detector, discontinued for low accuracy
OpenAI launched its own AI Text Classifier in January 2023 and withdrew it in July 2023, citing a “low rate of accuracy.” In its own reported testing, the classifier correctly identified only about 26% of AI-written text as “likely AI-written,” while incorrectly labeling human-written text as AI-generated about 9% of the time. That a detector’s own developer discontinued it on accuracy grounds is a useful benchmark for how far vendor-claimed accuracy can diverge from real-world performance, even for the organization that trained the underlying language model.
More recent peer-reviewed evaluation
Later evaluations published in the peer-reviewed International Journal for Educational Integrity (Springer) and in the journal Information (MDPI) have continued to assess commercial detectors, including Turnitin and Originality.ai, specifically in academic-integrity contexts, and report ongoing reliability limitations — inconsistent results across tools on the same text and documented false-positive and false-negative rates that the studies conclude make current detectors unsuitable as sole evidence in a misconduct determination. Because detector vendors update their underlying models on their own schedule, treat any specific percentage, including the figures above, as a snapshot of the version tested at the time, not a permanent characteristic of the product.
Institutions responding to reliability concerns
Reliability concerns have led at least some institutions to change how, or whether, they use AI-writing detection at all — Vanderbilt University is widely reported to have disabled Turnitin’s AI-writing detection feature specifically because of false-positive concerns. CASRAI’s comparison of Turnitin’s built-in AI detection against standalone detectors covers how the access model and reported accuracy evidence differ between institutionally licensed and directly purchased tools.
Why false positives cluster in predictable ways
The false-positive problem is not evenly distributed across all writing. The features that make text look statistically “AI-like” — low sentence-length variation, high word-choice predictability — are also features of writing that has been taught, templated, or heavily revised:
- Non-native English writers, per the Liang et al. findings above, are disproportionately flagged because second-language academic writing often relies on a narrower, more standard vocabulary and more templated sentence construction.
- Writers following a strict style guide or structured template — lab reports, structured abstracts, methods sections written to a journal’s required format — are constrained toward the same low-variance pattern by the genre itself.
- Heavily edited or professionally copyedited work can lose the idiosyncratic phrasing and sentence-length variation that reads as “human” to a detector, even when every word was authored or approved by the named writer.
For an integrity office, this means the population most likely to be false-flagged is not random: it skews toward international students, disciplines with rigid formatting conventions, and writers who received editing support — which is itself a due-process and equity consideration when a policy weighs how much evidentiary weight a flag should carry.
The “not sole evidence” principle in institutional guidance
Across the detector vendors, the independent research above, and the institutional policy examples CASRAI has documented in How Universities Are Updating Academic Integrity Policy for AI Writing Tools, a consistent position has emerged: an AI-detection score should be treated as one input that may prompt further inquiry, not as conclusive proof of misconduct on its own. Turnitin’s own guidance to institutions similarly cautions against treating a score as definitive. This mirrors long-standing practice for text-similarity reports, where a high similarity index has never, on its own, constituted a plagiarism finding without human review of the matched passages — see CASRAI’s iThenticate and Anti-Plagiarism Software resources for that parallel.
A practical evidentiary framework for integrity offices
The research above translates into a few concrete practices that a growing number of institutional policies now reflect:
- Treat a flag as the start of a review, not its conclusion. A detector score establishes reasonable cause to look further; it does not, by itself, establish that misconduct occurred.
- Require corroborating evidence before a finding. Draft history, version-control or track-changes logs, outlines, and correspondence with an advisor are process evidence a detector cannot see and cannot fabricate as easily as prose can be edited.
- Give the student or author an opportunity to respond before any finding is made, consistent with standard due-process practice in institutional disciplinary proceedings generally.
- Avoid a fixed numeric threshold presented as a bright-line rule. No detector vendor, and no field-wide standard body, has established a percentage above which a document is definitively AI-generated; a locally adopted cutoff is a policy choice, not a scientific finding, and should be documented and applied consistently rather than left to case-by-case discretion.
- Weight false-positive risk by population and context. Given the documented skew toward non-native English writers and formulaic academic genres, an investigation relying more heavily on a raw score for those cases carries more equity risk than the same score would in other contexts.
- Ask which report you’re looking at. A similarity report and an AI-writing report measure different things and require different follow-up — conflating them is a documented source of process error.
For guidance on writing these principles into an actual policy document — including disclosure requirements, tiered permission frameworks, and how to separate the data-privacy question from the integrity question — see CASRAI’s full guide to academic-integrity policy trends for AI writing tools. For the parallel question of how a flagged researcher or student can respond to a specific detection result, see Why Does My Paper Say “AI Detected”? and How to Detect AI-Generated Text in Academic Writing. For how this question differs from a plagiarism determination, see Is ChatGPT (AI-Generated Text) Considered Plagiarism?
Frequently asked questions
How accurate are AI-writing detectors like Turnitin, GPTZero, and Originality.ai?
Vendor-reported accuracy figures (for example, Turnitin’s claimed ~98% accuracy with under 1% false positives on documents where more than 20% of the text is AI-generated) describe internal testing under specific conditions. Independent peer-reviewed research has found substantially higher error rates in some populations — notably a 2023 Stanford study (Liang et al., published in Patterns) that found a 61.3% average false-positive rate across seven detectors on essays by non-native English speakers. Treat any single accuracy figure as conditional on the test conditions it was measured under, not a universal rate.
Can an AI-detection score alone be the basis for an academic-integrity finding?
Institutional guidance increasingly says no. Detector vendors themselves caution against treating a score as definitive, independent research has documented systematic false positives in specific writer populations, and a number of universities’ policies now explicitly require corroborating evidence — such as draft history or a conversation with the author — before a finding is made.
Why do non-native English speakers get flagged more often?
Detectors score text on predictability and sentence-length variation. Second-language academic writing often uses a narrower, more standardized vocabulary and more templated sentence construction, which produces the same statistical signature detectors associate with AI-generated text. This is a documented, published finding, not an anecdotal concern.
Have any universities stopped using AI-writing detection because of accuracy concerns?
Vanderbilt University is widely reported to have disabled Turnitin’s AI-writing detection feature specifically over false-positive concerns. Institutional practice varies; check a specific institution’s current policy rather than assuming a uniform approach.
Is there an industry-standard AI-detection accuracy threshold that counts as proof of misconduct?
No. No detector vendor and no field-wide academic-integrity standards body has published a numeric threshold that constitutes proof on its own. Any specific percentage cutoff in use at a given institution is a local policy choice, not an established scientific or field-wide standard.
Note on sourcing: figures in this guide are drawn from Turnitin’s own published claims, the peer-reviewed Liang et al. (2023, Patterns) study, OpenAI’s own public statement on discontinuing its AI Text Classifier, and peer-reviewed evaluations in the International Journal for Educational Integrity and Information. Detector vendors update models and accuracy claims on their own schedule; re-check current vendor documentation and recent literature before relying on any specific figure for a live case.







