Written and maintained by CASRAI Editorial Board
Last updated
Most people searching for the Copyleaks AI detection tool are not curious about the technology. They are deciding whether to buy it, renew it, or defend a finding it produced. This page is written for that decision, and specifically for the part of it that causes the most institutional pain: the false-positive rate.
The short version: Copyleaks is bought for breadth — LMS integration, multilingual coverage, and one screening pass across every submission in a term. It is not bought because it has the lowest false-positive rate, and the false-positive rate is precisely what an integrity office has to defend when a student challenges a finding. Those are two different jobs, and conflating them is how offices end up in unwinnable hearings.
Tip: try code CASRAI at checkout for 15% off, if the offer is currently active for this program — codes vary by vendor and aren’t guaranteed.
Where a second, accuracy-first opinion is genuinely useful, we point to Pangram — but as a targeted check on one disputed document, not as a replacement for institution-wide screening. We explain below exactly when that is the wrong call.
This page is deliberately Copyleaks-specific. If you are running a procurement exercise across several vendors, our four-tool comparison for integrity offices is the better starting point: AI detectors compared for research integrity offices. We will not re-litigate that here.
Copyleaks vs. Pangram at a glance
Copyleaks is bought for breadth — LMS-embedded screening across every submission and wide multilingual coverage. It is not bought for the lowest false-positive rate, which is what an integrity office must defend when a finding is challenged. An accuracy-first tool such as Pangram fits as a second opinion on one disputed document, not as an institution-wide replacement. No detector output is sole evidence of misconduct.
| Dimension | Copyleaks | Pangram |
|---|---|---|
| What it is bought for | Institution-wide screening embedded in the LMS, plus similarity checking and wide language coverage. | Accuracy-first checking of a specific document, typically as a second opinion after another tool has flagged it. |
| Accuracy claim (as of August 2026) | Not stated here. copyleaks.com returns HTTP 403 to automated retrieval, so no current vendor figure could be verified. Request it in writing during procurement, with the test set described. | Vendor self-reports 99.98% accuracy and roughly 1 false positive in 10,000 on aggregated public datasets. Vendor marketing, not an independent audit. |
| LMS integrations | Widely documented LTI integrations including Canvas, Moodle, Blackboard and D2L Brightspace. Confirm the current list with the vendor. | Educational Institution licence lists Canvas, Brightspace, Moodle and Google Classroom (vendor-verified August 2026). |
| Plagiarism / similarity matching | Mature similarity product line across web, published and internal sources. | Plagiarism detection included from the $20/month Individual tier upward (vendor-verified August 2026). |
| Cross-institutional student-paper repository | Repository behaviour depends on institutional configuration — confirm with the vendor. | Not described in Pangram published materials as of August 2026. If this is a hard requirement, evaluate Turnitin. |
| Multilingual coverage | Reported by third-party aggregators as substantially broader than accuracy-first competitors; not vendor-verified here. Test on your own languages. | Narrower. Weakest fit where non-English submission volume is high. |
| Free tier | Not verified. | 2,000 words per day, no payment method required (vendor-verified August 2026). |
| Published entry pricing | Quoted, not published. Institutional pricing is negotiated. | $20/month Individual (300,000 words/month); $65/month Professional (1,500,000 words/month, $200 API usage included); $20/seat/month Team, two-seat minimum. |
| API for batch screening | Available; pricing quoted rather than published, and not verified here. | Published per-word rates: Pangram 4 at $0.05 per 100 words, Pangram 3 at $0.05 per 1,000 words, with a bulk discount (August 2026). |
| Right role in an integrity workflow | The screening layer. Every submission, inside the gradebook, with an audit trail. | The corroboration layer. One disputed document, after a flag, before anyone is accused — most valuable when it disagrees. |
| Who should not switch | Nobody screening thousands of submissions a term in the LMS should leave this for a second-opinion tool. | Not a replacement for institution-wide screening, multilingual similarity checking, or an existing institutional contract. |
What a Copyleaks AI score actually measures (and what it cannot tell you)
A Copyleaks AI score is a classifier output. The model has been trained to separate text that looks statistically like machine generation from text that looks statistically like human writing, and it reports a probability-shaped number over a passage. That is all it is.
Three things follow from that, and all three matter in a hearing:
- It is a statement about textual features, not about authorship. The detector has no access to draft history, no access to the student, and no knowledge of who typed what. It observes the finished string and nothing else.
- It has no notion of permitted use. Most institutions now allow some AI assistance — outlining, grammar repair, translation support. A detector cannot distinguish disallowed generation from disclosed, permitted assistance, because the textual signature is similar or identical.
- A percentage is not a confidence interval. A “98% AI” label does not mean there is a 98% chance the student cheated. It is a model score on a scale the vendor chose, and the relationship between that number and real-world likelihood depends entirely on how common AI use actually is in the population being screened.
That last point is the one that most often goes wrong. Our guide on how AI detection actually works covers the underlying mechanics in more depth, and why does my paper say AI detected covers the same ground from the student’s side, which is worth reading before you sit across from one.
Copyleaks false-positive rate against published independent benchmarks
We are not going to give you a number here, and we want to be explicit about why.
First, a disclosure about method. In preparing this page we attempted to verify Copyleaks’ current published accuracy and false-positive claims directly against the vendor’s own live page. copyleaks.com returns HTTP 403 to automated retrieval, so we could not confirm any current vendor figure. Any accuracy percentage you see attributed to Copyleaks on a comparison site — including numbers repeated confidently across affiliate roundups — is either lifted from vendor marketing or from a third-party test whose methodology you have not seen. We are not going to launder one of those into a figure that looks verified because it appears on a nonprofit standards site. If you need Copyleaks’ current claimed false-positive rate, get it from the vendor in writing, as part of procurement, and get the test set described.
Second, and more importantly: published detector benchmarks generally do not transfer to your population. This is not a dodge, it is the actual methodological problem. A false-positive rate is only meaningful relative to the corpus it was measured on. Vendor and third-party benchmarks are typically run on aggregated public datasets — news text, Wikipedia extracts, essay corpora, model outputs from whichever generation of models was current. Your submissions are not that. They are discipline-specific, written to a prompt, by a cohort with a particular language profile, often after a writing-centre intervention, and increasingly with permitted AI assistance baked in.
Third, base rates dominate. If AI use in a given assignment is genuinely rare, even a very low false-positive rate produces a large share of false accusations among the papers that get flagged. This is ordinary conditional probability, and it is the single most consequential fact about deploying detection at scale. Screening ten thousand submissions with a detector that is wrong one time in a thousand still generates ten flagged innocents. Our guide on AI detection accuracy in higher education works through this arithmetic properly.
So the honest framing is not “which vendor has the lowest published false-positive rate.” It is “what is my process when the detector is wrong, because on volume it will be.”
See what a second-opinion check costs →
Why non-native-English and heavily-edited writing gets flagged more often
This is the best-documented failure mode in the field, and it is not specific to Copyleaks. Detectors keying on low lexical diversity and low sentence-level variation systematically over-flag writing that is fluent but linguistically conservative.
The landmark result is Liang, Yuksekgonul, Mao, Wu and Zou, “GPT detectors are biased against non-native English writers” (Patterns, 2023), which found that a range of contemporary detectors misclassified genuine TOEFL essays written by non-native speakers as AI-generated at strikingly high rates, while correctly classifying essays by native speakers. The mechanism is straightforward: a writer working in a second language tends to reach for a narrower, safer vocabulary and more regular sentence construction, which is exactly the signature detectors were trained to treat as machine-like.
The same logic catches several other groups:
- Heavily edited text. A paper that has been through a writing centre, a professional editor, or repeated self-revision converges toward conventional phrasing — the edits sand off precisely the idiosyncrasies that mark text as human.
- Formulaic academic genres. Methods sections, structured abstracts and systematic-review protocols are deliberately templated. Low variation is the disciplinary requirement, not a red flag.
- Writers using assistive technology. Grammar and style tools, dictation, and translation aids all push text toward the same conservative centre.
- Students with certain disabilities whose accommodations involve exactly those assistive tools.
The equity consequence is that a detector-led process, applied uniformly, does not fall uniformly. It concentrates on international students, on students who used the support services the institution told them to use, and on students with accommodations. An integrity office that cannot articulate this in a hearing is going to lose the hearing, and should.
LMS integration: Canvas, Moodle and where the check happens in the workflow
This is the real reason most institutions hold a Copyleaks licence, and it is a legitimate one.
Copyleaks is widely documented as offering LTI-based integrations with the major learning management systems — Canvas, Moodle, Blackboard and D2L Brightspace among them — so that the check runs as part of submission and the result lands where the marker already works. Confirm the current supported list and LTI version with the vendor during procurement rather than trusting any third-party list, including this one; integration catalogues change quietly.
The workflow placement is the point. When detection is embedded at submission:
- Every submission is screened, so there is no selection bias in who gets checked — a genuinely important fairness property.
- The marker sees the signal in context, alongside the similarity report and the submission itself.
- There is an audit trail tied to the gradebook entry.
Nothing about a standalone second-opinion tool replicates that, and no institution should try to run term-wide screening by having staff paste documents into a web form. If your requirement is “every submission, every course, inside the gradebook,” that requirement points at Copyleaks or Turnitin, full stop. Our comparison of Turnitin AI detection versus standalone AI detectors covers that trade in detail.
Licensing unit and pricing model at institutional scale
We could not verify Copyleaks’ current institutional pricing (see the 403 note above), and institutional detector pricing is negotiated rather than listed in any case. What we can tell you is which questions decide the cost, because the licensing unit varies between vendors and it is where budgets get surprised:
- Per credit, per page, per word, or per seat? A credit-based model priced per scanned page behaves very differently from a per-FTE site licence when a department decides to screen a backlog.
- Do re-scans consume credits? Resubmissions and appeals generate re-scans, and this is a common source of overrun.
- Is AI detection bundled with similarity checking, or separately metered?
- What happens at the cap — hard stop, overage billing, or silent degradation?
- Data retention and training. Ask, in writing, whether submissions are retained and whether they are used to train models. This is a procurement requirement at most institutions now, not a nice-to-have.
For contrast, and because these figures we could verify: as of August 2026, Pangram’s own published pricing page lists a free tier at 2,000 words per day with no payment method required; an Individual plan at $20/month covering up to 300,000 words per month with plagiarism detection included; a Professional plan at $65/month covering up to 1,500,000 words per month and including $200 in monthly API usage; a Team plan at $20 per seat per month with a two-seat minimum; and a separately-quoted Educational Institution licence described as unlimited checks with LMS integrations. Verify current figures against the vendor before budgeting — pricing pages move.
API access and batch screening a backlog of submissions
Batch screening a backlog is a genuinely different task from term-time screening, and it is where an API matters. Typical triggers: a policy change that applies retroactively, a journal or programme auditing a back catalogue, or an integrity office asked to review a cohort after a single confirmed case.
Both vendors expose APIs. Pangram publishes per-word API pricing directly — as of August 2026, its pricing page lists Pangram 4 at $0.05 per 100 words and Pangram 3 at $0.05 per 1,000 words, with a bulk discount — which makes a backlog job straightforward to cost out in advance. Copyleaks API pricing is quoted rather than published, and we could not verify it.
Two cautions before anyone runs a backlog job:
- Decide the disposition rule before you run the scan, not after. A retroactive sweep that produces two hundred flags with no pre-agreed threshold, no appeal route and no resourcing is an institutional crisis, not a finding. This is the most common way batch screening goes wrong.
- Check the retroactivity of your own policy. Screening 2023 submissions against a 2026 policy is usually indefensible, and detectors trained on current model outputs perform differently on older text anyway.
Try Pangram free — 2,000 words a day →
Using a second detector as corroboration rather than confirmation
Here is the distinction that does the work, and it is worth being pedantic about it.
Confirmation is running a second detector, getting a second flag, and treating two flags as stronger evidence than one. This is close to worthless. Detectors share training data, share architectural assumptions, and fail on the same inputs — the non-native-speaker bias above shows up across the field, not in one product. Two correlated instruments agreeing tells you much less than it feels like it does, and if the case is going to a hearing, “two tools agreed” invites the obvious question of whether they agreed for the same wrong reason.
Corroboration is different. It means using a second, independent signal to reduce the chance you are acting on a false positive — most valuably by looking for disagreement. A second detector that comes back clean on a document your primary tool flagged at 95% is genuinely informative: it tells you the flag is fragile and should not carry a case on its own. That is the use we think is defensible, and it is a use that mostly protects students rather than catching them.
Used that way, the second tool sits at one specific point in the workflow: a single disputed document, after a flag, before anyone is accused. That is a handful of checks per term, not a screening programme. It is also why we are comfortable recommending Pangram for this narrow role, and not comfortable recommending it as a Copyleaks replacement.
For what it is worth on the accuracy question: as of August 2026, Pangram’s own site claims 99.98% accuracy and a false-positive rate of roughly 1 in 10,000 on aggregated public datasets, with per-model figures of 99.8% against Claude Opus 5 and 99.6% against GPT 5.6. Those are vendor self-reported figures on the vendor’s chosen test data, not an independent audit — and per the base-rate discussion above, they are not a prediction of how the tool will behave on your students’ writing. Treat them the same way you should treat any vendor’s numbers, including Copyleaks’. Our guides on the most accurate AI detector and whether GPTZero is accurate apply the same scepticism across the field.
What Copyleaks does that Pangram does not — and who should stay with Copyleaks
We need to correct something that circulates widely in affiliate write-ups, because we checked it and it is wrong. Pangram is frequently described as having no LMS integration and no plagiarism checking. As of August 2026 that is not accurate: Pangram’s published pricing page lists an Educational Institution licence with integrations in Canvas, Brightspace, Moodle and Google Classroom, and plagiarism detection is included from its $20/month Individual tier upward. We are flagging this because we would otherwise be repeating a false differentiator in our own favour.
The real distinctions, stated carefully:
- Incumbency and workflow depth. Copyleaks is an established institutional product with years of deployment inside integrity workflows, existing contracts, existing staff training, and existing appeal precedent at your institution. That is a substantial, unglamorous advantage and it is usually the deciding one.
- Multilingual breadth. Copyleaks is consistently reported by third-party aggregators as covering a much wider set of languages for both AI and similarity detection than accuracy-first competitors. We could not vendor-verify the current count. If you screen substantial non-English submission volume, this is likely the single strongest argument for staying, and you should test it on your own languages rather than trusting a number.
- Similarity matching depth and repository behaviour. Copyleaks’ similarity checking against web, published and internal sources is a mature product line. Pangram’s published materials do not describe a cross-institutional student-paper repository of the kind Turnitin maintains; if matching against prior student submissions across institutions is a requirement, that points at Turnitin specifically rather than at either of these.
- Scale economics. At term-wide volume across a whole institution, a negotiated site licence from an incumbent will usually beat per-word or per-seat pricing from a smaller vendor. Run the arithmetic on your actual volume.
So, plainly: who should not buy Pangram? If you need detection embedded in the grading workflow across thousands of submissions every term, with mature multilingual similarity checking and an existing institutional contract, stay with Copyleaks — or evaluate Turnitin if cross-institutional repository matching is the requirement. A second-opinion accuracy tool does not solve that problem and buying one instead would be a mistake. Pangram earns its place on a specific disputed document, not across a catalogue.
If you are a teacher rather than an institutional buyer, our guide to the best AI detector for teachers and our standalone Pangram guide are more directly useful than this page.
What we verified, and what we could not
In the interest of letting you weigh this page properly:
- Verified against the vendor’s own live pages on 26 August 2026: all Pangram pricing, tier limits, API rates, accuracy claims and LMS integration claims quoted above.
- Not verified: every Copyleaks figure. copyleaks.com returns HTTP 403 to automated retrieval, so its integration list and multilingual coverage above are attributed to third-party aggregators and widely-published documentation, and are hedged accordingly. No Copyleaks accuracy or false-positive percentage appears anywhere on this page, because we could not source one honestly.
- Not independently tested: CASRAI has not run its own benchmark of either tool. Nothing here is a test result from us.
Frequently asked questions
What does a Copyleaks AI score actually mean?
It is a classifier output describing how statistically similar a passage is to machine-generated text. It is not a statement about authorship, it cannot distinguish disallowed generation from permitted and disclosed AI assistance, and a “98% AI” label does not mean a 98% chance the student cheated.
What is Copyleaks’ false-positive rate?
We could not verify a current figure: copyleaks.com returns HTTP 403 to automated retrieval, and we will not repeat an unsourced number. More importantly, a published benchmark rate rarely transfers to your population, because false-positive rates depend on the corpus measured and on how common AI use actually is among the submissions you screen. Request the figure and its test set in writing during procurement.
Why do non-native English speakers get flagged more often?
Detectors key on low lexical diversity and low sentence-level variation, which is exactly the profile of fluent but linguistically conservative second-language writing. Liang et al., “GPT detectors are biased against non-native English writers” (Patterns, 2023), found contemporary detectors misclassified genuine TOEFL essays as AI-generated at high rates while classifying native-speaker essays correctly. Heavily edited text, templated Methods sections and assistive-technology users are affected by the same mechanism.
Does running a second AI detector make a finding stronger?
Generally no. Detectors share training data and failure modes, so two agreeing flags are correlated rather than independent evidence. A second detector is most valuable when it disagrees — a clean result on a flagged document tells you the flag is fragile and should not carry a case alone.
Does Pangram have LMS integration and plagiarism checking?
Yes, as of August 2026 — contrary to how it is often described. Its published pricing page lists an Educational Institution licence with Canvas, Brightspace, Moodle and Google Classroom integrations, and plagiarism detection is included from the $20/month Individual tier upward.
Should we replace Copyleaks with Pangram?
Usually not. If you need detection embedded in the grading workflow across thousands of submissions, mature multilingual similarity checking, and continuity with an existing institutional contract, stay with Copyleaks — or evaluate Turnitin if cross-institutional repository matching is required. Pangram fits a targeted second-opinion check on a specific document.
Can a detector result alone support a misconduct finding?
No. No detector output from any vendor at any score is sole evidence of misconduct. Treat a flag as a trigger to look, never as a finding. What carries a case is draft and version history, a conversation with the student, inconsistency with prior work, fabricated citations, and the student’s own account.
The rule that outranks every number on this page
No detector output, from any vendor, at any score, is sole evidence of misconduct. Not Copyleaks, not Pangram, not two of them agreeing.
A defensible process treats a detector flag as a trigger to look, never as a finding. What actually carries a case is the ordinary evidence integrity work has always relied on: draft and version history, a conversation with the student about their own argument, inconsistency with prior submitted work, fabricated or unverifiable citations, and the student’s own account. If those are absent and all you have is a percentage, you do not have a case, and proceeding as though you do exposes the institution as much as the student.
Write that into the policy before you buy any tool. It is the part that determines whether the tool helps you or hurts you.








