Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & Research SupplyReagents, PPE & instruments — chain-of-custody documented.Fast, traceable sourcing built for regulated research environments, from bench consumables to instrumentation.Shop lac.us CodeCASRAIlac.us

How AI Detection Actually Works — and Why It Gets It Wrong

AI-text detectors measure perplexity and burstiness, not authorship. Here is how that mechanism actually works, why it produces predictable false positives against non-native English writers and other groups, and what a defensible institutional process requires before treating a detector score as evidence.

Ask about How AI Detection Actually Works — and Why It Gets It Wrong

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Last verified: 16 August 2026. This page explains the statistical mechanism behind AI-text detectors — what they actually measure, why that mechanism produces predictable false positives, and what a defensible institutional process requires when a detector flags someone’s work. It is written for two audiences: someone who has been flagged and needs to understand what happened, and the person — an integrity officer, editor, instructor, or committee member — who has to weigh a detector score as evidence. This page does not cover how to evade or defeat detection; that is a separate query intent this site deliberately does not serve.

What an AI-detection tool actually measures

Every mainstream AI-text detector — Turnitin’s AI writing indicator, GPTZero, ZeroGPT, Originality.ai, Copyleaks — is built on the same underlying idea, regardless of the specific model or marketing language each vendor uses. None of them “recognize” AI text the way antivirus software recognizes a known malware signature. Instead, they estimate two statistical properties of a passage:

  • Perplexity — how predictable each word is, given the words before it, according to a language model. Text where every word choice is close to what a language model would have predicted next has low perplexity. Large language models are, by construction, optimized to produce low-perplexity text: they are trained to output the statistically most likely continuation.
  • Burstiness — how much sentence length and structure vary across a passage. Human writing tends to alternate between short, punchy sentences and long, complex ones, and to vary rhythm unevenly across a document. Machine-generated text tends to be more uniform: sentences cluster around a similar length and structural pattern.

A detector runs a passage through a language model (often a smaller open-source model, not the exact model that may have generated the text), computes perplexity and burstiness scores, and compares them to a threshold or a trained classifier built on labeled human/AI examples. The output — a percentage, a label, a color-coded highlight — is a statistical estimate that a passage resembles the low-perplexity, low-burstiness profile of machine-generated text. It is not a determination of fact, and no vendor’s own documentation claims it is.

Why detectors cannot identify which model wrote something

A common misconception, on both sides of an allegation, is that a detector has somehow identified “ChatGPT” or “Gemini” as the source. It has not, and structurally cannot. Large language models are trained on overlapping corpora and share similar objectives, so their output distributions converge on similar statistical profiles — low perplexity, low burstiness — regardless of which company built the model. A detector measuring those two properties is measuring a signature that is common across most modern LLMs, not a fingerprint unique to one of them. Any tool or vendor claiming to name a specific model with confidence is overstating what the underlying method supports.

Why human writing gets flagged: the false-positive problem

Because perplexity and burstiness are proxies for “predictable, uniform text” rather than direct evidence of machine authorship, any human writing that happens to be unusually predictable or uniform trips the same signal. This is not a rare edge case; it is a structural weakness in the method, and it clusters in identifiable kinds of writing:

  • Non-native English writing. This is the best-documented false-positive pattern. A peer-reviewed study by Liang and colleagues, published in the journal Patterns (Cell Press) in 2023, tested seven widely used GPT detectors against a set of TOEFL essays written by non-native English speakers and against essays written by native-English-speaking US eighth-graders. The detectors produced an average false-positive rate of 61.3% on the non-native-writer essays — more than three in five were flagged as AI-generated — against a near-zero false-positive rate on the native-speaker essays. The researchers attributed this to non-native writers tending to use simpler, more predictable vocabulary and sentence construction, which drives perplexity down for entirely human reasons.
  • Formulaic or heavily templated academic prose. Structured sections such as methods descriptions, literature-review summaries, or standard IMRaD-format writing follow conventional phrasing and predictable structure by design — that predictability is often a mark of good, disciplined technical writing, not machine authorship.
  • Heavily edited or proofread text. Multiple editing passes tend to smooth out the unevenness that drives burstiness up. A student, author, or editor who revises a draft repeatedly for clarity is, as a side effect, making the text look statistically more like a low-burstiness machine output.
  • Technical, legal, or scientific writing generally. Domains with constrained vocabulary and conventional sentence structures produce naturally low-perplexity text regardless of who wrote it.

None of these patterns describe rare writers. They describe a large share of the population a university, journal, or funder actually deals with — international students and scholars, technical writers, careful editors — which is why detector scores cannot be treated as free-standing evidence.

Why detection degrades on short passages

Perplexity and burstiness are statistical estimates computed over a sample of text, and like any statistical estimate, their reliability depends on sample size. A single sentence or a short paragraph provides too little data for the underlying model to distinguish signal from noise, so detector confidence intervals widen sharply below roughly a few hundred words. This is why a single flagged paragraph pulled out of a longer, otherwise-unflagged document is a particularly weak basis for a finding: the tool is operating well outside the range where it has any meaningful discriminating power, even before accounting for the false-positive patterns above.

Is ZeroGPT accurate?

ZeroGPT is one of the widely used free detectors and uses the same general perplexity/burstiness-based approach as the rest of the category, marketed under its own “DeepAnalyse” branding. ZeroGPT’s own site describes a “high accuracy model” trained across multiple languages and LLM outputs, but — consistent with most detector vendors — does not publish an independently audited accuracy or false-positive rate. No detector vendor’s self-reported accuracy claim should be treated as equivalent to independent, peer-reviewed testing: vendor-reported numbers are typically measured against the vendor’s own test set, under conditions the vendor controls, and are not directly comparable across tools. Because ZeroGPT relies on the same statistical mechanism described above, there is no methodological reason to expect it to be exempt from the false-positive patterns documented for other detectors in that category — non-native English writing, formulaic prose, and short passages remain the same predictable risk factors. For a side-by-side look at how specific detectors compare, see AI Detectors for Research-Integrity Offices and Is GPTZero Accurate?, which examines the same question for a specific, independently studied tool.

Why no detector can prove authorship

A detector score is a probability estimate about statistical resemblance, not a forensic determination. There is no watermarking standard in general use across major LLM providers that would let a third-party tool cryptographically verify that a specific passage came from a specific model (or from no model at all), and no detector vendor claims their tool provides that. That means a detector score, on its own, can suggest a pattern worth asking about — it cannot establish, on its own, that a specific person did or did not write a specific passage. Treating a score as proof inverts what the tool is actually capable of measuring.

If you have been falsely accused of using AI

If a detector has flagged your work, the strongest response is not to argue with the score in the abstract but to produce independent evidence of your own authorship and writing process:

  • Draft and version history. Cloud word processors (Google Docs’ version history, Microsoft Word’s track changes, or a document’s file-system modification history) record incremental edits over time in a way that is very difficult to fabricate after the fact. This is usually the single strongest piece of corroborating evidence available to someone falsely accused.
  • Research notes, outlines, and source material. Drafts, annotated sources, citation manager libraries, and notes that predate the final document support a genuine research and writing process.
  • The institution’s actual written policy. Ask specifically what the policy says about how a detector score may and may not be used as evidence, and what your right of reply is. Policies vary significantly between institutions and between publishers, and most formal academic-integrity and publication-ethics processes give the accused person a defined opportunity to respond before any finding is made.
  • Request the underlying evidence, not just the score. A percentage alone is not an explanation. Ask what specific passages were flagged and why, so you can address the actual basis for the concern rather than the headline number.

For the process specific to a Turnitin AI-writing flag, see Turnitin AI Detection: How It Works and What a Score Actually Means.

What a defensible institutional process requires

For the person or committee on the other side of an allegation, the same mechanism limits what a detector score can responsibly support. A defensible process treats a detector flag as a trigger for further inquiry, not as a finding in itself, and generally requires:

  • Corroborating evidence beyond the score — version history, prior drafts, interview or discussion with the accused, consistency with the person’s documented prior work, and any other evidence that bears on authorship independent of the detector’s statistical output.
  • A documented, published policy that specifies how detector output may be used, what threshold (if any) triggers review, and who makes the final determination — rather than an ad hoc reaction to a single number.
  • A right of reply that gives the accused a genuine opportunity to present evidence and explanation before a finding is recorded, consistent with the general due-process expectations that apply to research-misconduct and academic-integrity findings more broadly.
  • Awareness of the false-positive population described above — non-native English speakers and international scholars in particular are structurally more likely to be flagged by chance alone, which has real equity implications for how an institution weighs a flag against those populations.
  • Human review by someone trained on the tool’s actual limitations, not an automated pass/fail gate. Several institutions have scaled back or discontinued standalone AI-detector use specifically because of false-positive risk; see AI Detector False-Positive Controversies for recent examples.

These principles track the general evidentiary standard used across research-integrity work: findings of misconduct are expected to rest on a preponderance of evidence assembled through an inquiry-and-investigation process, not on a single automated signal. For the fuller evidentiary framework aimed specifically at academic-integrity offices, see AI Detection Accuracy in Higher Education, and for how a multi-signal investigation is actually structured, see How to Detect AI-Generated Text in Academic Writing. For how ORI findings of research misconduct work more generally, see Research Misconduct: What an ORI Finding Actually Means.

Frequently asked questions

How does AI detection work for essays specifically?

The mechanism is identical to any other text: the tool computes perplexity and burstiness across the essay and compares the result to a threshold or trained classifier. Essays are a common flagpoint specifically because student essays are often short, follow assignment-driven templates, and are written by a population that includes many non-native English speakers — all three of which are documented false-positive drivers, independent of whether the essay was actually AI-written.

What causes an AI detection false positive?

The main documented causes are non-native English writing patterns, formulaic or template-driven prose, heavily edited or proofread text, technical/legal/scientific writing with constrained vocabulary, and passages too short for the statistical estimate to be reliable. Each of these independently produces the same low-perplexity, low-burstiness statistical profile the detector is trained to associate with machine-generated text.

I’ve been falsely accused of using AI. What should I do?

Gather version history from your word processor, prior drafts, research notes, and any source material that predates the final document. Ask for the institution’s written policy on how detector scores may be used as evidence, and ask what specific passages were flagged and why. Most formal integrity processes provide a right of reply before any finding is made — use it to present this evidence rather than only disputing the score itself.

How do I prove I didn’t use AI?

There is no single document that “proves” human authorship the way a receipt proves a purchase, but incremental version history (Google Docs’ version history or Word’s track-changes/file history) is the closest available equivalent, because it shows a document being built up over time through edits that are difficult to fabricate retroactively. Combined with research notes, outline drafts, and consistency with your other documented writing, this constitutes the strongest practical evidence available.

Is ZeroGPT accurate?

ZeroGPT does not publish an independently audited accuracy or false-positive rate, and self-reported vendor accuracy claims are not directly comparable to peer-reviewed testing. Because it uses the same perplexity/burstiness-based approach as other detectors in this category, it is reasonable to expect the same documented false-positive risk factors — non-native English writing, formulaic prose, short passages — to apply to it as well.

Can AI detectors tell you which AI model was used?

No. Detectors measure statistical properties (perplexity and burstiness) that are broadly shared across most modern large language models, not a signature unique to one model or vendor. A tool or report that names a specific model with confidence is going beyond what the underlying method can actually support.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →