Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & Research SupplyReagents, PPE & instruments — chain-of-custody documented.Fast, traceable sourcing built for regulated research environments, from bench consumables to instrumentation.Shop lac.us CodeCASRAIlac.us

The most accurate AI detector: how to actually judge the claim

Every AI detector claims 99% accuracy. Here is what that number hides, which error rate matters for integrity cases, and how the main tools compare.

Ask about The most accurate AI detector: how to actually judge the claim

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Best evidence · Verified 18 August 2026

Pangram — the lowest published false positive rate, independently evaluated

Free tier · Individual $20/mo

On the metric that actually matters for academic integrity — how often human writing gets wrongly flagged — Pangram publishes the strongest claim in the category (roughly 1 in 10,000) and points to third-party evaluation by University of Chicago and University of Maryland researchers rather than only its own benchmarks. Independent evaluation is uncommon here, and on a false-positive claim it is the difference between evidence and marketing. Free tier: 2,000 words a day.

Try Pangram free Opens on the vendor’s site · CASRAI referral link

Different use case? Compare Originality.ai and GPTZero → — Accuracy is not the only axis. Volume auditing and classroom use pull towards different products regardless of detection performance.

Editorial disclosure: CASRAI has commercial referral arrangements with some of the vendors named on this page, and may earn a commission if you subscribe to them. We name them here regardless of whether a link is present. We only recommend tools our editorial team has independently researched. Read our full disclosure policy.

In summary

  • Headline “accuracy” combines two very different errors. Ask for the false positive rate specifically — it is the one that harms students.
  • At 10,000 submissions a term, a 1% false positive rate wrongly flags 100 students. At 0.01% it flags one.
  • Vendor benchmarks are self-selected. Third-party evaluation, where it exists, is worth far more than a bigger number.
  • Detector performance degrades on edited, translated or paraphrased text, and shifts whenever underlying language models change.
  • No detector should be the sole basis for an integrity finding. Test any candidate against your own known-human documents before deploying.

The detectors worth shortlisting

Ranked for institutional academic-integrity and editorial use rather than for casual checking. Every one of these should be tested against your own documents before deployment — a vendor benchmark tells you about their test set, not yours.

#1 Pangram — our winner

The strongest false-positive evidence, and the only one pointing at independent evaluation.

Best for: Integrity offices and editorial teams who must defend a result to a student, author or appeals panel

Price: Free tier · Individual $20/mo

Claims 99.98% accuracy and a roughly 1-in-10,000 false positive rate, with evaluation by University of Chicago and University of Maryland researchers. LMS integrations cover Canvas, Moodle, Google Classroom and Brightspace, with an API for editorial workflows. The free tier at 2,000 words a day is enough to run a real evaluation against your own back-catalogue before spending anything.

Strengths

  • Lowest published false positive rate in the category
  • Points to third-party evaluation rather than only vendor benchmarks
  • Broad LMS coverage plus an API for editorial screening
  • Genuinely usable free tier for evaluation

Trade-offs

  • Newer entrant with a shorter track record than incumbents
  • Independent evaluation is a snapshot — model landscape shifts constantly
  • Institutional licensing is custom-quoted rather than transparent

Try Pangram free CASRAI referral link — disclosed on this page

#2 Originality.ai

Built for auditing content at volume, with plagiarism and readability bundled in.

Best for: Editorial offices and integrity teams screening large numbers of submissions

Price: From $14.95/mo (2,000 credits)

Credit-based pricing suits bursty, high-volume screening, and the bundled plagiarism, readability and fact-check tooling means one subscription rather than three. Unlimited team seats with per-member activity logs matter when several people share the screening workload and you need to know who checked what.

Strengths

  • Designed for volume auditing rather than one-off checks
  • Plagiarism checking bundled into every plan
  • Unlimited team seats with activity logging

Trade-offs

  • Accuracy figures are vendor-reported
  • No LMS integration — a real constraint for classroom deployment
  • Credit model needs monitoring so you do not run dry mid-term

Try Originality.ai CASRAI referral link — disclosed on this page

#3 GPTZero

The one actually designed around the classroom conversation.

Best for: Instructors and academic-integrity offices needing LMS integration

Price: Free tier · Individual from ~$9.99/mo

Sentence-level highlighting is the differentiator, and it is a pedagogical feature rather than a technical one: it gives an instructor something specific to discuss with a student instead of a single opaque percentage. LMS integration across Canvas, Moodle and Blackboard means it appears where marking already happens, which is what determines whether a detector gets used consistently.

Strengths

  • Sentence-level highlighting supports a real conversation
  • LMS integration where marking already happens
  • Free tier available

Trade-offs

  • Reported accuracy varies widely between studies (roughly 87–99.5%)
  • Less suited to high-volume editorial screening

Try GPTZero free CASRAI referral link — disclosed on this page

#4 QuillBot

Free, convenient, and fine as a first look — not an institutional tool.

Best for: Individual writers sanity-checking their own work before submission

Price: Free tier · Premium from ~$8.33/mo billed annually

Bundled into a suite most students already use, which makes it the detector people actually reach for. That convenience is the point and also the limit: it is a reasonable first look for an individual, and it is not built for institutional deployment, appeals defensibility or volume screening.

Strengths

  • Free and immediately available
  • Sits alongside tools writers already have open

Trade-offs

  • Not designed for institutional or appeals use
  • No LMS integration or audit trail
  • Published false-positive evidence is thin

Try QuillBot free CASRAI referral link — disclosed on this page

What a 99% accuracy claim actually hides

Accuracy is the proportion of classifications a detector gets right across a test set. That single figure hides everything a buyer needs to know, for three reasons.

It combines two non-equivalent errors. Missing AI-generated text and wrongly flagging human text are both “inaccuracy”, and they have wildly different consequences. A single accuracy figure lets a vendor trade one against the other — a detector tuned to flag aggressively catches more AI and also accuses more innocent people, and its accuracy number can stay high throughout.

It depends entirely on the test set. A detector evaluated against unedited ChatGPT output versus polished academic prose will look superb. The same detector evaluated against lightly-edited AI text, translated text, or text from a fluent non-native English writer will look considerably worse. Vendors choose their own test sets and rarely publish the composition, which makes cross-vendor comparison of accuracy figures close to meaningless.

It depends on base rate. If 5% of submissions are AI-generated, a detector could achieve 95% accuracy by flagging nothing at all. Accuracy figures quoted without the class balance of the test set cannot be interpreted.

The practical consequence: stop asking “which is most accurate” and start asking “what is your false positive rate, on what test set, evaluated by whom”. A vendor that can answer all three is telling you something. One that answers only the first is not.

The arithmetic institutions should run

Take an institution processing 10,000 submissions in a term, and assume — generously to the detector — that 5% are genuinely AI-generated. That is 500 AI submissions and 9,500 human ones.

A detector with a 1% false positive rate wrongly flags 95 of the human submissions. If it catches 90% of the AI ones, it correctly flags 450. So of 545 total flags, roughly one in six is a false accusation. Any process treating a flag as strong evidence is going to produce a substantial number of wrongful outcomes.

Now the same institution with a detector at a 0.01% false positive rate: one wrongly-flagged human submission against 450 correct ones. Now a flag genuinely is meaningful evidence.

The detection rate barely moved the outcome; the false positive rate transformed it. This is why the ordering of questions matters, and it is why an institution should insist on a false positive figure before comparing anything else. It is also why the free-tier evaluation described below is not optional — a vendor’s FPR on their test set is not your FPR on your students’ writing.

Why detection performance degrades

Even a detector that performs well today will perform differently in six months, for reasons that are structural rather than fixable.

Editing destroys the signal. Detectors work on statistical properties of text — word choice predictability, sentence structure regularity, perplexity and burstiness. A person who takes AI output and rewrites a third of it disrupts exactly those properties. Detection performance on heavily-edited text is much worse than on raw output, and heavily-edited output is what a motivated student actually submits.

Paraphrasing tools disrupt it deliberately. This is an acknowledged limitation across the category, and it is why any claim of near-perfect detection should be read as applying to unmodified text.

New models shift the distribution. Detectors are trained on the output of models that exist at training time. Each new generation writes differently, and detectors need retraining to keep up. There is an inherent lag, and it favours the generator.

The blurred middle is growing. The binary framing — human or AI — increasingly does not describe reality. A researcher who drafts their own argument and uses a tool to tighten the prose has produced something that is genuinely both, and a detector has no way to represent that. This is why disclosure-based policies are, in our view, more durable than detection-based ones: they ask what the writer did, which is the actual question, rather than inferring it from textual statistics.

How to evaluate a detector properly

Do not buy on a published figure. Run your own evaluation — it takes an afternoon and it is the only thing that tells you how a tool behaves on your population.

  1. Assemble known-human documents from your own context. Work submitted before generative AI was widely available is ideal, because provenance is certain. Aim for at least a hundred pieces.
  2. Deliberately include writing by non-native English speakers in proportion to your actual student body. This is the population most at risk of false flags, and a detector that performs badly here will produce discriminatory outcomes however good its aggregate figure looks.
  3. Include assistive-technology users and highly formulaic writers where you can identify them, for the same reason.
  4. Run the batch and count the false positives. That number, on your documents, is the only one that matters.
  5. Then test detection with AI-generated text at realistic levels of editing — not raw output, but output someone has genuinely worked over.
  6. Test the workflow, not just the engine. If markers must copy text into a separate tool, it will be used inconsistently and abandoned within a term. Test the LMS path.
  7. Re-run it annually. Performance shifts as models change; a one-off evaluation ages badly.

Free tiers make this genuinely cheap: Pangram gives 2,000 words a day, GPTZero and QuillBot both have free tiers. An institution can evaluate three detectors properly for nothing before committing to any of them.

Detection is not a policy

The institutions that handle this well are the ones that decided what they wanted students to do before deciding how to check. Detection is a control; it is not a position.

A workable policy states what use of generative AI is permitted for which kinds of assessment, requires disclosure of what was used and how, explains what happens when a concern is raised, and gives students a genuine route to respond. Within that structure a detector is a triage instrument that decides where to look more closely — a supporting tool, doing a defined job.

Without that structure, a detector becomes the policy by default: markers make individual judgements from a percentage, decisions vary by department, appeals succeed inconsistently, and the students most harmed are the ones least able to contest a process nobody has written down. Buying the most accurate detector does not fix that. It just makes the wrong decisions more confidently.

Run the evaluation on your own documents

The free tier gives 2,000 words a day — enough to work through a batch of known-human writing from your own institution, including non-native English speakers, and get the one false-positive number your appeals process will actually care about.

Free tier · Individual $20/mo

Try Pangram free Opens on the vendor’s site · CASRAI referral link

Frequently asked questions

Which AI detector is the most accurate?

On the metric that matters for academic integrity — false positive rate — Pangram publishes the strongest claim, at roughly 1 in 10,000, and points to third-party evaluation by University of Chicago and University of Maryland researchers. But no published figure substitutes for testing a detector against your own known-human documents, because performance depends heavily on the writing population.

What does 99% accuracy actually mean for an AI detector?

Less than it sounds. Accuracy combines missed AI text and wrongly-flagged human text into one figure, so a detector tuned to flag aggressively can keep a high accuracy number while accusing more innocent people. It also depends entirely on the test set the vendor chose and on the proportion of AI text within it, neither of which is usually published.

How many students would a 1% false positive rate wrongly flag?

At 10,000 submissions in a term, 100. If roughly 5% of submissions are genuinely AI-generated and the detector catches 90% of them, that produces around 545 flags of which about 95 are false — meaning roughly one in six flags is a wrongful accusation.

Do AI detectors work on edited or paraphrased text?

Much less well. Detectors rely on statistical properties of text such as word predictability and sentence regularity, and editing disrupts exactly those properties. Paraphrasing tools disrupt them deliberately. Any near-perfect detection claim should be read as applying to unmodified model output rather than to what a motivated student submits.

Can we use an AI detector result to fail a student?

Not on its own, and no reputable position supports it. A detection score is a probabilistic signal about textual features, not evidence of what a student did. It should open an enquiry in which other evidence is gathered — draft history, version records, a discussion of the content — with the student given a real opportunity to respond.

How should we evaluate a detector before buying?

Assemble at least a hundred known-human documents from your own context, deliberately including writing by non-native English speakers, run them through the free tier, and count the false positives. Then test detection against realistically-edited AI text rather than raw output. Test the LMS workflow too, and re-run the whole evaluation annually.

Is detection or disclosure the better approach?

Disclosure is more durable. Detection infers what a writer did from statistical properties of the finished text, which breaks down as the middle ground between human and AI writing grows. Disclosure asks the actual question. A detector is a useful triage instrument within a disclosure-based policy, and a poor substitute for one.

Related on CASRAI

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →