Skip to main content
v2026.11,610 entries · CC-BY 4.0

Editorial · CASRAI · AI and ML research outputs

A Bad Dataset Came Down for Copyright, Not for Being Wrong

Kaggle removed a flawed stroke-detection dataset in July 2026 after a copyright complaint — five months after a research-integrity complaint about the same data went nowhere. Three Scientific Reports papers built on it have now been retracted.

Published 6 Aug 2026· 7 minute read

Ask about this story

Answers are drawn from this article and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

CASRAI is the reference for research administration — bookmark it for the next question.

In late July 2026, Kaggle removed a stroke-detection image dataset that had already underpinned a retracted Scientific Reports paper, continued to be used in at least two papers published earlier in 2026, and had been the subject of a research-integrity complaint since February. None of that integrity concern got it removed. What worked, five months later, was a copyright complaint from a dental journal.

The gap between those two outcomes is the actual story here, and it matters more to anyone building or reusing research datasets than the specifics of one bad Kaggle upload do.

What was in the dataset

The dataset in question was a Kaggle collection titled “Facial droop and facial paralysis image,” stored internally in a folder named “droopy,” and marketed as a source of stroke-patient facial images for training machine-learning models. A Scientific Reports paper published in December 2025 used it to train a predictive model for early stroke detection.

Adrian Barnett, a medical statistician at Queensland University of Technology in Australia, and Alexander Gibson, a PhD student working with him, first documented problems with this and related Kaggle datasets in a medRxiv preprint posted in February 2026, which Nature‘s news team covered in April. Using reverse image search, Barnett matched four photographs in the dataset to images published years earlier in the Journal of the Canadian Dental Association (JCDA). The people in those photographs had Bell’s palsy, not stroke. Retraction Watch’s reporting on the audit also documented, among the dataset’s other images, celebrity and stock photography — including repeated images of well-known actors — mixed in with the ostensibly clinical material, images of children included without any documented consent, and no verifiable record of confirmed diagnoses or ethical oversight for any of the images. In short: a training set marketed as clinical stroke data that was, in substantial part, not clinical data at all.

Barnett flagged the problems to Kaggle, to the journal, and to the paper’s authors starting in February 2026. Kaggle’s public response at the time, as reported by Retraction Watch, was that its datasets are intended for “benchmarking and development” rather than as primary evidence for medical research, and that the platform did not consider the dataset to violate its terms of service. The dataset stayed up.

What changed things was the JCDA itself contacting Kaggle in July 2026 to assert that its copyrighted images had been used without permission. That complaint is what Kaggle acted on, removing the dataset for intellectual-property infringement rather than for the underlying data-integrity problems Barnett had raised five months earlier. For a research-integrity audience, that enforcement asymmetry is the headline: a copyright holder with a specific, ownable claim got a faster and more decisive response than a methodologist raising a broader, harder-to-adjudicate concern about whether the data was fit for its stated clinical purpose at all.

The retractions so far, and what’s still open

As of the July 2026 reporting, the original December 2025 Scientific Reports stroke-detection paper had been retracted, with a retraction date of June 29, 2026, and two further Scientific Reports papers relying on datasets flagged in the same audit had also been retracted for lacking verifiable data provenance. Springer Nature, which publishes Scientific Reports, had identified eleven flagged papers across its journals in total, with several more under active investigation; Tim Kersjes, the publisher’s head of research integrity, has described the response as ongoing and case-by-case rather than a single blanket action. Elsevier and MDPI journals were also carrying flagged papers — nine and eleven respectively, per Barnett and Gibson’s count — with investigations announced but not concluded at either publisher as of this reporting.

Not every author has treated the underlying data as the problem. Naeem Ramzan, an author of one flagged paper, defended the dataset’s use by pointing to roughly two dozen other published articles that also relied on it — a response that, read alongside Barnett’s findings, illustrates rather than rebuts the concern: wide reuse of an unverified dataset spreads the same provenance problem across the literature rather than validating it. Alaa Mohamed, the corresponding author of another flagged paper at Mansoura University, did not respond to Retraction Watch’s requests for comment. Separately, Ben Van Calster, a biostatistician at KU Leuven, has said he has observed comparable data-provenance problems in COVID-19 prediction models, suggesting this is not confined to one dataset or one clinical application.

Barnett and Gibson’s broader audit traced 124 published papers back to the flagged Kaggle datasets, across multiple publishers, and found those 124 papers had in turn been cited by roughly 86 review articles — a pattern Barnett has termed a “laundering effect,” in which an unverified dataset’s problems become progressively harder to see as they move through secondary literature. Perhaps most notably for anyone tracking whether flagged data actually stops getting used: Retraction Watch reported that researchers have continued to use the dataset even after the concerns became public, identifying at least two papers published in 2026 that relied on it, including one July 2026 preprint that defended the choice on the grounds that the dataset is “widely adopted” in the machine-learning literature. Wide adoption is not the same claim as verified provenance, and treating the two as interchangeable is precisely the gap this episode exposes.

What comes next

Barnett’s team is now planning a considerably larger follow-on project: a quality audit of 100 datasets hosted on Kaggle, assessing whether they are fit for use in medical research. The stated aim is to establish whether the “droopy” dataset’s problems — misattributed provenance, misclassified subject matter, no ethical documentation, no confirmed diagnoses — are an isolated failure or a pattern across the platform’s medical-adjacent datasets more broadly. Given that the original audit alone surfaced 124 downstream papers from a handful of flagged datasets, a 100-dataset review is likely to generate a sustained pipeline of further findings rather than a single conclusive result.

Why this matters beyond one dataset

For research data management practice, the practical lesson isn’t really about Kaggle specifically — it’s about what “provenance” needs to mean before a third-party dataset is treated as evidence, not just as training material. A few things this case makes concrete:

  • Platform terms of service are not a fitness-for-purpose signal. Kaggle’s own position, before removal, was that the dataset was fine for “benchmarking and development” but was never represented as validated clinical data — a distinction the papers built on top of it did not preserve.
  • “Widely used” is not “verified.” Two dozen other papers relying on the same dataset, or a defense that a dataset is “widely adopted,” describes exposure, not provenance. Citation count and reuse frequency tell you nothing about whether the underlying images or records are what they claim to be.
  • Reverse image search and manual provenance checks found what automated checks and prior peer review did not. The Bell’s palsy misclassification, the JCDA image reuse, and the celebrity photos were all identifiable by a domain expert doing manual verification — not flagged by any dataset-quality gate before publication.
  • Enforcement follows the clearest legal claim, not the most serious problem. A copyright holder’s narrow, provable claim moved faster than a methodologist’s broader integrity concern. Anyone relying on a platform’s complaint process to catch bad research data should not assume integrity concerns alone will be sufficient, or fast.
  • Downstream literature obscures upstream problems. With 124 papers built on the flagged datasets and roughly 86 reviews citing those papers, the “laundering effect” means a dataset’s original defects can be several citation-hops removed from anyone encountering the claim it ultimately supports.

None of this is new territory for research data management as a discipline — documenting source, custody, and transformation of a dataset before treating it as reusable evidence is exactly what provenance practice and standards such as PROV-O exist to formalize, and what a documented data management plan is meant to make explicit before data collection even begins. What this case adds is a concrete, dated illustration of what happens when that documentation is simply absent from a dataset that nonetheless gets treated, downstream, as adequate for clinical machine-learning research.

Further reading on CASRAI

Sources: Retraction Watch, “Kaggle removes problematic stroke dataset for copyright infringement” (July 29, 2026); Retraction Watch, “‘Comically bad’ datasets used to train clinical models for stroke and diabetes” (May 18, 2026).

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →