Verdict · Verified 18 August 2026
Pangram — the strongest false-positive evidence in the category
Free tier · Individual $20/mo
Pangram claims 99.98% accuracy and a false positive rate of roughly 1 in 10,000, and — more importantly — points to third-party evaluation by researchers at the University of Chicago and the University of Maryland rather than only its own benchmarks. Independent evaluation is rare in this category, and on a false-positive claim it is the difference between a marketing figure and evidence. The free tier gives 2,000 words a day, which is enough to test it against your own real documents before spending anything.
Try Pangram free Opens on the vendor’s site · CASRAI referral link
Compare it against Originality.ai and GPTZero → — No detector, including this one, should be the sole basis for an academic-integrity decision. That position is not a hedge — see the section on what a detection score can and cannot support.
Editorial disclosure: CASRAI has commercial referral arrangements with some of the vendors named on this page, and may earn a commission if you subscribe to them. We name them here regardless of whether a link is present. We only recommend tools our editorial team has independently researched. Read our full disclosure policy.
In summary
- Pangram claims 99.98% accuracy and a false positive rate of about 1 in 10,000, with third-party evaluation by University of Chicago and University of Maryland researchers.
- Pricing: Free (2,000 words/day), Individual $20/mo (300,000 words), Team $20/seat/mo (min 2 seats), Professional $65/mo (1.5M words plus $200 API credit). Verified 18 August 2026.
- Integrations include a web dashboard, Chrome extension, Google Docs, an API, and LMS support for Canvas, Moodle, Google Classroom and Brightspace.
- False positive rate matters far more than headline accuracy at institutional scale — at 10,000 submissions a term, a 1% FPR means 100 wrongly-flagged students.
- A detection score is evidence to open a conversation, never a finding on its own. Build your process around that or the tool will cause harm.
Pangram plans
Verified from the vendor pricing page, 18 August 2026
| Dimension | Free | Individual | Team | Professional |
|---|---|---|---|---|
| Price | $0 | $20/mo | $20/seat/mo (min 2) | $65/mo |
| Word allowance | 2,000/day (~60,000/mo) | 300,000/mo | 300,000/mo per seat | 1,500,000/mo |
| Image scans | 3/day | 100/mo | 100/mo per seat | 500/mo |
| API credit included | No | No | No | $200/mo |
| Annual saving | — | $60/yr | $60/seat/yr | $240/yr |
A separate Developer API is priced in tiers from $25 to $1,000, with a stated 20% bulk discount. Institutional licences are custom-quoted — which is the route an integrity office or editorial department should take rather than stacking individual seats.
Why false positive rate beats accuracy
“99% accurate” sounds decisive and tells you almost nothing useful, because it collapses two very different errors into one figure. A detector makes two kinds of mistake: it can miss AI-generated text (a false negative), or it can flag human-written text as AI (a false positive). Those are not equivalent, and in an academic-integrity context they are not remotely equivalent.
A false negative means someone gets away with it. That is bad. A false positive means a student who wrote their own essay is accused of cheating, has to prove a negative, and may face a process that damages their record, their visa status or their mental health. The asymmetry is enormous, and it is why the false positive rate is the number an integrity office should ask for first.
Now apply scale. Imagine an institution processing 10,000 submissions in a term:
- At a 1% false positive rate — which would be an unremarkable figure for many detectors — that is 100 wrongly-flagged students in one term.
- At 0.1%, ten students.
- At Pangram’s claimed 1 in 10,000 (0.01%), one.
The difference between those scenarios is not a technical nicety. It is the difference between a workable process and one that generates a hundred wrongful accusations a term, most of which will land hardest on the students least equipped to contest them.
Pangram’s published comparison claims false positive rates of 0–0.8% across document types against competitors at 0.5–2.4%. Treat any vendor’s own comparison with appropriate scepticism — but the underlying point stands regardless of whose numbers you use: ask every vendor for their false positive rate on text resembling your actual student population, and treat a refusal to give one as an answer in itself.
Third-party evaluation is rare here
The AI detection market has a structural evidence problem: nearly every accuracy claim comes from the vendor, tested against a dataset the vendor chose, using a methodology the vendor did not publish. That is not evidence in any sense a researcher would accept.
Pangram points to evaluation by researchers at the University of Chicago and the University of Maryland, which puts it in a small group. This matters because independent evaluators choose their own test sets, and the composition of the test set is where detection benchmarks are most easily flattered — a detector evaluated only on unedited ChatGPT output will look far better than one evaluated on text that has been edited, translated, or written by someone whose English is highly fluent but non-native.
Two caveats worth holding. Independent evaluation is a snapshot: models change constantly, and a detector’s performance against a 2026 model says little about its performance against next year’s. And “researchers at institution X evaluated it” is not the same as “institution X endorses it” — read the actual study rather than the citation, and check what was tested and against what.
Still, the direction of travel is right. A vendor willing to be evaluated by people it does not control is making a different kind of claim from one that is not.
What a detection score can and cannot support
This is the section that matters most, and it applies to Pangram exactly as it applies to every competitor.
A detection score is a probabilistic signal, not a finding of fact. It is a reason to look more closely at a piece of work. It is not, by itself, evidence that a student cheated, and no institution should operate a process in which a score alone produces a sanction. If your academic-integrity procedure allows that, the procedure is the problem, not the tool.
Some groups are structurally more likely to be flagged. Research has repeatedly found that text by non-native English writers is disproportionately flagged by AI detectors, because the features detectors associate with machine generation — limited vocabulary variation, conventional sentence structures, predictable phrasing — overlap with the features of competent second-language academic writing. Autistic students, students using assistive writing technology, and students taught to write in a highly formulaic register can be affected similarly. An institution deploying detection without accounting for this will produce discriminatory outcomes, whatever the aggregate false positive rate says.
A confident-looking percentage invites over-reading. “98% likely AI-generated” reads as near-certainty to a marker under time pressure. It is a model output about textual features, not a measurement of what a student did.
What a defensible process looks like: a flag opens an enquiry rather than a case; the student is told what has been observed and given a genuine opportunity to respond; other evidence is gathered — draft history, version records, viva-style discussion of the content, comparison with previously submitted work; the detection score is one input among several and is never the sole basis for a finding; the student can see and contest the evidence; and the whole thing is documented consistently enough to survive an appeal.
Used that way, a detector with a genuinely low false positive rate is a useful triage instrument. Used as an oracle, it will produce injustices at a rate proportional to your submission volume.
Where it plugs in
Pangram offers a web dashboard for direct text and file upload, a Chrome extension, a Google Docs integration, and LMS integrations for Canvas, Moodle, Google Classroom and Brightspace. There is also an API for institutional and developer use, and an image detector in research preview.
The LMS integration is the one that decides whether this works at institutional scale. A detector requiring markers to copy text into a separate web tool will be used inconsistently and abandoned within a term; one that surfaces alongside the submission gets used. If you are evaluating for institutional deployment, test the LMS path specifically rather than the dashboard.
The API matters for a different buyer — journal editorial offices screening submissions at volume, where the sensible pattern is integration into an existing manuscript workflow rather than a manual step. The Professional tier includes $200 of monthly API usage, which is a reasonable starting point for evaluating that path before committing to a Developer API tier.
On the Google Docs integration specifically: it is useful for the draft-history conversation described above, because it lets you look at how a document developed rather than only at the finished text. Draft history is often more informative than any detection score.
Who should and should not buy this
Worth evaluating if you are an academic-integrity office that has to defend detection results to students and appeals panels; a journal editorial office screening submissions where a wrongful accusation against an author carries real reputational cost; or an institution that has already deployed a detector and is seeing a false-positive problem you cannot defend.
Probably not worth it if your institution has not yet settled its policy on what a detection result may be used for. Buying a detector before agreeing the process is the wrong order, and it reliably produces the situation where markers make ad-hoc decisions with no institutional backing. Settle the policy first; the tool is the easy part.
Definitely not if you intend to use scores as automatic findings. In that configuration a better detector produces more confident wrong decisions, not fewer wrong ones.
Practical evaluation advice: use the free tier’s 2,000 words a day to run your own known-human documents through it — ideally including work by non-native English writers in your actual student population — before running anything else. A vendor’s benchmark tells you how it performs on their test set. Your own back-catalogue tells you how it will perform on yours, and that is the only number your appeals process will care about.
Test it against your own documents first
The free tier gives 2,000 words a day — enough to run a batch of known-human work, including writing by non-native English speakers, and see the false positive behaviour on text that actually resembles your population.
Free tier · Individual $20/mo
Try Pangram free Opens on the vendor’s site · CASRAI referral link
Frequently asked questions
How accurate is Pangram?
Pangram claims 99.98% accuracy with a false positive rate of roughly 1 in 10,000, and points to third-party evaluation by researchers at the University of Chicago and the University of Maryland. Independent evaluation is rare in this category, though it is a snapshot — detector performance shifts as underlying language models change.
How much does Pangram cost?
Free at 2,000 words a day, Individual at $20/month for 300,000 words, Team at $20 per seat per month with a two-seat minimum, and Professional at $65/month for 1.5 million words plus $200 of API credit. A Developer API runs from $25, and institutional licences are custom-quoted. Verified 18 August 2026.
Why does false positive rate matter more than accuracy?
Because the two errors are not equivalent. A false negative means someone gets away with it; a false positive means an innocent student is accused and must prove a negative. At 10,000 submissions a term, a 1% false positive rate produces 100 wrongly-flagged students, against one at a 0.01% rate. That difference decides whether a process is workable or harmful.
Can a Pangram result be used as proof a student used AI?
No, and no detector’s result should be. A detection score is a probabilistic signal about textual features, not a finding of fact about what a student did. It is a reason to open an enquiry and gather other evidence — draft history, version records, a discussion of the content — never a basis for a sanction on its own.
Do AI detectors discriminate against non-native English speakers?
Research has repeatedly found that text by non-native English writers is disproportionately flagged, because the features detectors associate with machine generation overlap with those of competent second-language academic writing. Students using assistive writing technology and those taught a highly formulaic register can be affected similarly. Any institutional deployment needs to account for this explicitly.
Does Pangram integrate with our LMS?
Yes — Canvas, Moodle, Google Classroom and Brightspace, alongside a web dashboard, Chrome extension, Google Docs integration and an API. If you are evaluating for institutional use, test the LMS path rather than the dashboard: a detector that requires copying text into a separate tool gets used inconsistently and abandoned.
Should we buy a detector before setting our AI policy?
No. Buying the tool before agreeing what a detection result may be used for reliably produces ad-hoc decisions by individual markers with no institutional backing, and those are the cases that fail on appeal. Settle the policy and the process first — the tool is the easy part.







