Examples
Worked examples
- Is an instance
A review-article manuscript at 28% similarity where the overlap is entirely correctly quoted and cited excerpts plus a shared reference list -- not treated as a concern once the report is opened and read.
- Is an instance
An institution's internal graduate-school policy that flags any electronic thesis deposit above 20% similarity for a supervisor's manual review before acceptance -- an informal local trigger, not a field-wide rule.
Counter-examples
Looks similar, but isn't
- Not an instance
A manuscript at 6% similarity where the small flagged portion is a closely paraphrased, uncited paragraph concentrated in the paper's own Discussion section -- a low percentage that does not resolve the concern, since the tool cannot detect uncredited paraphrasing on its own.
Editorial commentary
A similarity or originality report — generated by Turnitin Similarity, Crossref’s Similarity Check, or iThenticate — returns a percentage showing how much of a submitted manuscript, thesis, or dissertation textually matches material already in the tool’s reference corpus. Authors, students, and even some editors often ask what percentage counts as “acceptable.” The honest answer is that no universal number exists: no publisher body, funder, or the U.S. Office of Research Integrity’s federal misconduct definition (42 CFR Part 93, which is scoped to fabrication, falsification, and plagiarism as concepts, not to a text-matching score) sets a percentage threshold. Any specific number in circulation is a local, informal policy, not a field-wide standard.
Common informal benchmarks — and why they aren’t standards
Individual universities, graduate schools, and some journals do publish internal guidance suggesting a similarity index above roughly 15% or 20% should trigger a closer look before submission or deposit is accepted. These figures are useful as an internal triage trip-wire — a signal to open the report and read what’s underlined, not skip straight to it — but they vary widely between institutions, are set locally rather than by any standards body, and are not binding outside the institution that published them. A manuscript submitted to one journal at 22% similarity might draw no comment; the same report at a different venue might prompt an editor’s query. Treating either number as a pass/fail line misunderstands what the tool measures.
Why the percentage alone is a blunt instrument
A similarity index is a measure of textual overlap, not a measure of misconduct, and several ordinary, entirely legitimate things inflate it:
- Properly quoted and cited material — block quotations with quotation marks and a citation still register as a textual match.
- Reference lists and bibliographies — shared citation formatting and repeated journal/publisher names across a field’s literature match against thousands of other papers citing the same sources.
- Common methodological or boilerplate phrasing — standard language in a Methods section (e.g., a widely used assay protocol description) or a clinical trial’s model agreement boilerplate is shared across many papers by design.
- A candidate’s own prior work — a thesis chapter that reuses text from the candidate’s own previously deposited coursework or preprint inflates the score without involving anyone else’s material (see self-plagiarism for how that case is handled differently from third-party plagiarism).
None of these indicate misconduct on their own. Conversely, a low score does not clear a paper: text-matching tools only catch verbatim or lightly reworded overlap, so they miss uncredited paraphrased ideas, fabricated or falsified data, and text deliberately reworded through automated paraphrasing tools — the latter sometimes leaves a detectable trace known as tortured phrases, where an unusual synonym substitution lowers the similarity score while making the underlying appropriation just as improper.
What matters instead: what’s actually flagged
Turnitin’s own guidance to institutions is explicit that its Similarity tool “does not determine whether plagiarism has occurred” — it is built to support a reviewer’s judgment, not substitute for it. In practice that means an editor, supervisor, or integrity officer needs to open the report and check: is the highlighted text quoted and cited correctly? Is it concentrated in a few sources or spread thin across many (a pattern more consistent with common phrasing than copying)? Is any of it in the paper’s own original argument or results, where overlap would be far harder to justify than in a literature-review paragraph summarizing prior work? The Committee on Publication Ethics (COPE) guidance on plagiarism and text overlap is consistent with this: any resulting action follows a case-by-case editorial or institutional investigation, never an automatic decision keyed to the percentage alone.
Worked examples
Example 1 — high score, no misconduct. A systematic review manuscript returns a 28% similarity index. On inspection, the report shows the overlap is almost entirely block-quoted excerpts from the studies under review (each correctly quoted and cited) plus a lengthy shared reference list with other reviews on the same topic. An editor reasonably concludes the score reflects the nature of a review article, not appropriation.
Example 2 — low score, real concern. A manuscript returns a 6% similarity index — comfortably under any informal institutional benchmark — but the small amount flagged is concentrated entirely in two consecutive paragraphs of the paper’s own Discussion section, closely paraphrasing a single uncited source with only minor wording changes. The low overall percentage does not resolve the concern; the concentration and location of the match is what an editor would investigate.
Related terms
See Originality Report for how the similarity index itself is calculated and which tools produce it, Self-Plagiarism for how a candidate’s or author’s own prior text is treated differently from third-party matches, Plagiarism for the underlying misconduct concept, and Tortured Phrases for the forensic signature of text reworded specifically to evade a similarity check.
References
- Turnitin Guides, “Understanding the similarity score”
- Crossref, “Similarity Check” service documentation
- Committee on Publication Ethics (COPE), Core Practices and case guidance on plagiarism and text overlap
- U.S. Office of Research Integrity, 42 CFR Part 93 (definitions of research misconduct)
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="Acceptable Similarity Percentage"
vocab-term-identifier="https://casrai.org/dictionary/term/acceptable-similarity-percentage" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/acceptable-similarity-percentage",
"name": "Acceptable Similarity Percentage",
"identifier": "https://casrai.org/dictionary/term/acceptable-similarity-percentage",
"description": "There is no single, official percentage that marks a similarity/originality report as acceptable or unacceptable. Turnitin, Crossref Similarity Check, and iThenticate all publish a numeric similarity index, but none of the vendors, COPE, or ORI define a universal cutoff -- individual institutions, journals, and graduate schools sometimes set an informal internal benchmark (commonly cited figures cluster around 15% or 20%) purely as a triage trigger for closer human review, not as a formal misconduct threshold. What actually determines whether a score is a problem is what the flagged text consists of: properly quoted and cited material, a shared reference list or bibliography, standard field terminology, and (for theses) a candidate's own previously deposited coursework all inflate the percentage without indicating misconduct, so a report has to be opened and read, not just scored, before any percentage is meaningful.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/research-integrity#set",
"url": "https://casrai.org/dictionary/term/acceptable-similarity-percentage",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"dateModified": "2026-07-23T07:22:15",
"inLanguage": "en"
}






