Through most of 2025, an “AI Detected” flag on a paper was a private hassle: a conversation with an instructor, a frustrating appeal, maybe a resubmission. In 2026, the same flag has become something bigger — the subject of a court ruling, and the reason a growing number of universities are turning the underlying detection feature off entirely. For researchers, graduate students, and the administrators who handle academic-integrity cases, the practical question has shifted from “how do these tools work” to “how much weight can this score legally and institutionally bear.” (For the mechanics of why detectors misfire on human writing in the first place, see CASRAI’s guide Why Does My Paper Say “AI Detected”? Understanding False Positives — this piece covers the 2026 developments building on top of that explanation.)
A courtroom test case: Newby v. Adelphi University
The clearest sign that AI-detection false positives have moved from an academic-integrity dispute to a legal-liability question is a New York case that concluded in early 2026. Orion Newby, an Adelphi University student who received tutoring support through the university’s Bridges program for students with learning and neurological disabilities, was accused of academic dishonesty after a professor ran his paper through an AI-detection tool. Newby maintained the writing was his own, produced with help from a university tutor. He was given a zero on the assignment, and a second alleged instance raised the possibility of expulsion.
Newby’s parents retained attorney Mark Lesko and sued. A New York state Supreme Court judge ruled in the family’s favor, reversing the disciplinary findings and ordering Adelphi to expunge the record and restore Newby to good standing. Lesko called the court’s opinion “groundbreaking,” telling reporters that “higher education needs to take a very careful look at this.” The case is notable less for any specific accuracy statistic and more for what the ruling establishes: that a student facing serious academic consequences is entitled to a meaningful due-process review before an AI-detection score alone can support a finding of misconduct.
One state-court ruling does not bind other jurisdictions or institutions, and it is not the last word on how much process is legally required elsewhere. But it is one of the first instances of a court, rather than an internal appeals committee, weighing in on an AI-detection dispute at all, and institutions and higher-education attorneys are treating it as a signal of the liability now attached to leaning on a single detector score.
Universities are turning the AI-writing check off
Separately from the litigation, a number of universities have begun disabling AI-writing detection features while keeping traditional text-matching (similarity/originality) checking in place — treating the two as genuinely different tools with different reliability records. Curtin University announced it will disable Turnitin’s AI-writing detection feature from January 1, 2026, while continuing to use the platform’s originality-checking functionality for academic-integrity purposes; the university framed the move around strengthening “trust and clarity” in assessment and keeping practices “fair and relevant.” Commenting on the decision, Dr Mark A. Bassett, an academic lead in artificial intelligence at Charles Sturt University, described Curtin as “joining the growing list of providers that are abandoning this deeply flawed technology” — a characterization from one commentator rather than an official industry tally, but one that reflects a real and growing current of institutional skepticism, not an isolated decision.
The pattern matters for researchers and graduate students because institutional policy on AI detection is genuinely in flux right now, in both directions — some institutions are tightening screening (see CASRAI’s guide to AI detection tools adoption in academic publishing for how publishers are simultaneously expanding editorial-integrity screening, a related but distinct mechanism from author-facing AI-writing detectors) while others are removing it. Assuming last year’s policy, or another institution’s policy, still applies is no longer a safe assumption.
What the underlying accuracy research still says
Both the litigation and the institutional reversals sit on top of accuracy findings that were already well documented before 2026 and remain the evidentiary backbone of both. The most frequently cited is a 2023 study by Liang, Yuksekgonul, Mao, Wu, and Zou (Stanford University), published in the journal Patterns, which tested seven GPT detectors against 91 TOEFL essays written by non-native English speakers and found an average false-positive rate of 61.3%, with more than 91% of the essays flagged by at least one detector — against a near-zero false-positive rate on essays by native-English-speaking U.S. eighth-graders. The same study found that light, prompt-based rewriting could push genuinely AI-generated text below detection thresholds, meaning the same tools were simultaneously over-flagging real human writing and under-flagging actual AI output.
Nothing about the 2026 developments changes that underlying picture; if anything, the Newby case and the institutional walk-backs are best read as courts and universities catching up to accuracy limitations researchers had already flagged. Detector vendors, including Turnitin, have themselves long cautioned against treating an AI-writing score as sole or definitive evidence — see CASRAI’s fuller explainer on why false positives happen and what to do if you’re flagged for the complete mechanics (perplexity, burstiness, and why formulaic or non-native academic writing is disproportionately affected) and a step-by-step response guide.
What this means for researchers and institutions right now
For individual researchers and graduate students, the practical guidance has not changed even though the institutional landscape has: keep draft history, version-controlled files, and correspondence with advisors as your strongest evidence of authentic authorship, since that is exactly the kind of process documentation a detector cannot see and a due-process review can. Disclose any legitimate, policy-permitted AI assistance rather than leaving it undisclosed — see CASRAI’s Generative-AI disclosure statement entry for what a compliant disclosure typically covers, and the Detection tool (AI-generated) and Generative AI dictionary entries for the underlying definitions. Text-matching similarity checks (Turnitin Similarity, iThenticate, Crossref Similarity Check) remain a separate, more mature technology from AI-writing detection — see CASRAI’s Anti-Plagiarism Software guide for how those compare.
For institutions and administrators, the throughline across both the Newby ruling and the Curtin decision is the same: an AI-writing percentage, on its own, is increasingly treated as an insufficient basis for a misconduct finding, and policies that don’t build in a meaningful human review and appeal step ahead of any consequence carry real legal and reputational exposure. Given how quickly both case law and institutional policy are moving in this area, any specific accuracy percentage, vendor claim, or list of participating institutions should be treated as a snapshot of a fast-moving situation rather than a settled fact — check your own institution’s or target journal’s current written policy directly rather than assuming it matches what’s described here.







