Written and maintained by CASRAI Editorial Board
Last updated
Short answer: GPTZero doesn’t flag “everything” as AI, but it does flag a real, documented, and specific class of human writing more often than it should — and there’s an actual statistical reason for it, not a conspiracy or a broken tool. If a piece of writing you know is human keeps coming back flagged, it’s very likely landing in one of a small number of well-understood trigger categories: non-native/multilingual English, technical or formulaic academic prose, template-driven answers, or short responses with too little text to score confidently. This guide explains the actual mechanism behind that pattern, who it hits hardest, what GPTZero itself says about it, and what an instructor or a flagged student should actually do about a disputed score.
Tip: try code CASRAI at checkout for 15% off, if the offer is currently active for this program — codes vary by vendor and aren’t guaranteed.
The short answer
- GPTZero scores text using signals tied to how predictable and how uniform the writing is — not whether a human actually wrote it.
- Human writing that happens to be unusually predictable and uniform (cautious ESL phrasing, rigid technical formatting, short boilerplate answers) scores the same way AI-generated text does, because the signal genuinely can’t tell the two apart in those cases.
- GPTZero has publicly acknowledged this and named the affected groups itself; it has also changed its underlying model since the “perplexity and burstiness” framing became widely known, though those concepts are still part of how it explains a result.
- A GPTZero score is a statistical estimate about a text’s writing pattern, not a determination about who wrote it. No detector output, from GPTZero or any competitor, is sufficient evidence on its own to conclude misconduct.
How GPTZero actually decides something looks like AI
The two concepts most people have heard in connection with GPTZero are perplexity and burstiness. In GPTZero’s own explanation, perplexity measures how likely it is that a language model would have chosen the exact words in front of it — predictable, high-probability word choices score low perplexity, while unexpected word choices score high. GPTZero has stated that, on its own scale, “a perplexity above 85 is more likely than not from a human source.” Burstiness measures how much that predictability varies across a document: large language models tend to hold a fairly constant level of predictability sentence to sentence, which GPTZero calls an “AI-print,” while human writers naturally swing between simple and complex sentence construction as they write.
Put simply: low perplexity plus low burstiness — consistently predictable wording, sentence after sentence — is the pattern GPTZero associates with machine-generated text. That is a real, coherent statistical signal. It is also a pattern that shows up in human writing whenever a person writes in a genuinely constrained way, for entirely human reasons that have nothing to do with authorship.
One nuance most explainers skip: GPTZero itself states that, as of autumn 2023, perplexity and burstiness are no longer the primary basis for its detection score — the company says it migrated to a deep-learning classifier trained directly on text patterns, with perplexity and burstiness retained as one of several supporting “indicators” used mainly to explain a result to the user, alongside things like sentence-level highlighting. The underlying vulnerability is the same either way: a trained classifier still learns to associate certain surface patterns — uniform structure, conventional phrasing, low lexical surprise — with AI output, because those patterns are genuinely more common in AI text on average. It just makes that association implicitly inside a neural network instead of computing it as an explicit formula. That doesn’t remove the bias; it makes the bias harder to audit from the outside.
Who actually gets hit hardest
This isn’t a vague complaint — GPTZero has published its own list of writer categories most prone to false positives, and it lines up closely with independent academic testing:
- Multilingual and non-native English (ESL) writers. Careful, conservative word choice — the exact thing ESL writers are often taught to do to write correctly — reads statistically like low perplexity.
- Technical and formulaic academic writing. Method sections, protocol descriptions, and other genres that reward precision and repetition over stylistic variety naturally compress toward the same low-perplexity, low-burstiness pattern.
- Template-based or heavily structured responses. Answers written to a rigid format (a rubric, a standard reporting structure, a fixed set of required headings) mimic the uniformity a detector associates with machine output.
- Short responses. There simply isn’t enough text for any statistical method to score confidently, so short answers are inherently noisier — in either direction.
The independent evidence for the ESL pattern specifically predates GPTZero’s current model but is the reason this concern is taken seriously at all: a 2023 Stanford-led study (Liang et al., Patterns, Cell Press) tested seven widely used AI-text detectors — including GPTZero — against 91 real TOEFL essays written by non-native English speakers and 88 essays written by native-English-speaking US eighth-graders. Across the seven detectors, the average false-positive rate on the TOEFL essays was 61%, versus near-perfect accuracy on the native-speaker essays from the same detectors; 89 of the 91 TOEFL essays were flagged as AI-generated by at least one detector, and all seven detectors unanimously flagged 18 of them. The study tied this directly to perplexity: the essays flagged unanimously had measurably lower perplexity than the rest, consistent with the mechanism above. When the same TOEFL essays were revised with more idiomatic, native-sounding word choices, the average false-positive rate dropped by roughly 49 percentage points.
GPTZero has responded to exactly this concern in its own public materials, stating that “our efforts in reducing ESL bias in classification since April 2022 have reduced AI detection’s false positive rate on TOEFL texts to 1.1%,” attributing the improvement to model parameter tuning, text preclassification, and adding more representative training data. Treat that 1.1% figure as a vendor-reported number describing GPTZero’s own later testing, not an independently replicated result on the same scale as the Stanford study — the two numbers aren’t measuring the same model version or necessarily the same test set, so the honest position is that GPTZero says it has substantially improved on this specific failure mode, and that claim has not been independently re-tested at the scale of the original study.
Check a document with GPTZero →
What a GPTZero score is — and isn’t
A GPTZero result is a statistical estimate about writing pattern, generated by a model with a known, acknowledged blind spot for certain categories of legitimately human text. It is not a lie-detector test, it is not a plagiarism finding, and by itself it does not establish that a specific person did or didn’t write a specific document. This is true of every AI-text detector on the market, not a weakness unique to GPTZero — see our companion breakdown of how AI detection actually works and why it gets it wrong for the mechanism across the category, and what independent accuracy testing on GPTZero specifically shows before treating a score as decisive either way.
That distinction matters most in exactly the moment a flag actually gets used: a disputed academic-integrity case. No credible academic-integrity policy should treat a detector score alone as sufficient evidence of misconduct, and most institutions that publish explicit AI-detection policies say so directly — a flag is a prompt to look further, not a finding.
What to do if your own writing gets flagged
None of the following is about evading detection — it’s about handling a false positive on writing you actually wrote:
- Keep your process evidence before you ever submit. Version history (Google Docs, Word’s track changes, a Git-tracked draft), notes, outlines, and earlier drafts are the strongest rebuttal to a flag, because they show the writing developing over time in a way a single generated pass wouldn’t.
- Read the highlighted sentences, not just the headline score. GPTZero highlights the specific passages driving the result. If they cluster in your most technical or most formulaic section, that’s consistent with the false-positive pattern above, not proof of anything.
- Ask what your institution’s policy actually requires. Most integrity processes require corroborating evidence beyond a detector score before any finding is made — know what that threshold is before you argue the score itself.
- Request a human conversation, not just a re-run. Re-scanning the same text with the same tool won’t resolve a mechanism-driven false positive; a conversation about your drafting process will.
What instructors and integrity offices should do with a disputed flag
Given the documented pattern above, a defensible process looks like this:
- Treat the score as one input, never the sole basis for a misconduct finding.
- Ask the student for process evidence (drafts, version history, notes) before escalating.
- Weigh whether the flagged text falls into one of the four high-false-positive categories — ESL, technical/formulaic writing, template-driven answers, short responses — before assuming the flag is meaningful.
- Document the full basis for any finding in writing, not just “GPTZero flagged this at X%.”
- Apply the same standard consistently across students — a policy that only gets scrutinized when a student objects isn’t a policy.
If your institution needs this written down formally, CASRAI’s reference entries on the generative AI disclosure statement and AI tool disclosure cover the adjacent question of what a student or author should proactively disclose, which is a cleaner starting point for policy than adjudicating detector scores after the fact.
Is GPTZero still worth using, given all this?
Here’s the honest tradeoff: GPTZero’s ESL and formulaic-writing false-positive bias is real, it’s documented by an independent peer-reviewed study, and GPTZero itself has publicly acknowledged it rather than denying it — which is unusual and, on balance, a point in its favor for transparency, but it doesn’t cancel out the underlying problem. That combination caps how strongly this page can recommend GPTZero as a purchase: it is a genuinely useful triage tool for surfacing text worth a closer look, and its willingness to publish its own limitations and offer sentence-level explanations is better practice than staying silent about the issue. It is not a tool that should be the last word on an integrity decision, for the reasons above, and no detector currently is.
Where GPTZero is a reasonable fit: a first-pass screening tool inside a broader review process that already requires corroborating evidence, used by someone who understands and applies the ESL/technical-writing caveat before escalating anything. Where it isn’t: as sole, automatic grounds for an academic-integrity finding, or as a tool aimed at non-native English speakers without that caveat built into the process around it. If your institution has already run into this problem and is shopping for a different fit, our GPTZero alternatives comparison covers how Pangram, Originality.ai, and Copyleaks handle the same false-positive pattern differently.
Frequently asked questions
Does GPTZero really flag “everything” as AI-generated?
No — but it does over-flag a specific, well-documented set of writing styles: non-native/ESL English, technical or formulaic academic prose, rigidly templated answers, and short responses. If your experience is “everything I write gets flagged,” it’s worth checking whether your writing consistently falls into one of those categories, since that’s the actual mechanism at work, not random malfunction.
Why did GPTZero flag text I know I wrote myself?
Most likely because the passage scored as unusually predictable and uniform in wording — the pattern GPTZero’s model associates with AI generation — for reasons that have nothing to do with who wrote it, such as careful ESL phrasing, technical repetition, or a short length that gives the model too little to work with.
Is a GPTZero score proof that someone used AI?
No. It’s a statistical estimate of writing pattern, not a determination of authorship. No credible academic-integrity process should treat a detector score alone as sufficient evidence of misconduct.
Does GPTZero still use perplexity and burstiness to detect AI text?
Not as the primary method. GPTZero states it moved to a deep-learning classifier as of autumn 2023; perplexity and burstiness are still surfaced as explanatory indicators for a result, but they are no longer described as the core detection signal.
Can a paraphrasing tool make GPTZero stop flagging my writing?
This page isn’t the place for detector-evasion advice, and using a tool specifically to defeat detection raises its own integrity questions. The better response to a false positive is process evidence — drafts and version history — and a human review, not trying to alter your writing until a score changes.
How accurate is GPTZero overall, not just on false positives?
Accuracy varies meaningfully by source text and use case; see our dedicated breakdown of what independent testing on GPTZero actually shows for the fuller picture beyond the false-positive question covered here.








