Verified against primary and near-primary sources on 9 July 2026. This is an actively moving area of law and policy — the summary below reflects the state of EU, UK, and US text-and-data-mining (TDM) and copyright law as of that date, and it will go stale. Re-check before relying on it for a compliance decision.
When a research project uses text, images, code, or other copyrighted material to train or fine-tune a machine learning model — or produces a model or dataset that others will train on — two separate questions arise: is the mining lawful (a copyright/TDM-exception question), and can anyone tell what was used (a provenance-documentation question). The two are connected but distinct, and research administration sits squarely on the second one: documenting where training data came from, under what rights, and how it was processed is a research data management (RDM) discipline, not just a legal afterthought bolted on at publication.
What “training data provenance” means operationally
CASRAI’s Dictionary already defines the underlying concept — see Training data provenance and the related AI provenance term for the operational definitions and worked examples. In short: provenance documentation for a training dataset records where each component of the data came from, under what licence or legal basis it was obtained, what processing was applied (filtering, deduplication, annotation), and what consent or rights-reservation status applies to it. This guide builds on that definition to cover the current legal landscape around lawful mining, and how to fold provenance documentation into standard RDM practice — the DMP, the metadata record, the data availability statement — rather than treating it as a separate exercise.
Why this is an RDM problem, not only a legal one
Research offices already have infrastructure for exactly this kind of documentation: the Data Management Plan, structured metadata schemas, and licence/reuse-permission tracking (see Creative Commons Licenses for Research Data and the reuse license and copyright dictionary terms). Training-data provenance is the same discipline applied to a new kind of research output: instead of only documenting the provenance of a measured or observed dataset, a project that trains, fine-tunes, or redistributes a model needs to document the provenance of the corpus that shaped it.
This matters for reasons that predate and outlast any specific court ruling:
- Reproducibility. A model’s behavior can’t be independently assessed or reproduced without knowing what it was trained on — the same logic that already drives data availability statements (see How to Write a Data Availability Statement for Reproducibility).
- Downstream compliance. If a training corpus or a resulting model is later shared, published, or commercialized, whoever redistributes it needs to know its rights basis — the same problem CASRAI’s guidance on choosing an open data repository already addresses for conventional datasets.
- Funder and institutional accountability. Machine-actionable DMPs (see machine-actionable DMP and the RDA DMP Common Standard) are built precisely so that structured metadata like this — sources, licences, processing steps — can be captured once and reused across the project lifecycle instead of reconstructed after the fact.
The legal landscape: TDM exceptions and AI training, jurisdiction by jurisdiction
None of this replaces institutional legal advice. What follows is a factual summary of where the law currently stands, sourced directly from the statutes, the EU AI Office, and primary court-ruling coverage.
European Union: DSM Directive Articles 3 and 4
Directive (EU) 2019/790 on Copyright in the Digital Single Market created two separate TDM exceptions:
- Article 3 is a mandatory exception for research organisations and cultural heritage institutions carrying out text and data mining for the purposes of scientific research, on works to which they already have lawful access. Rightsholders cannot contract this exception away.
- Article 4 is a broader exception available to anyone, for any purpose (including commercial AI training) — but rightsholders may reserve their rights, including by machine-readable means, and once reserved, the material falls outside the exception.
The Directive required a Commission review report no sooner than 7 June 2026 — a deadline that has just passed as of this guide’s verification date; no legislative amendment to Articles 3/4 has been made as a result of that review as of writing. Source: Directive (EU) 2019/790, Articles 3–4, via EUR-Lex.
EU AI Act: Article 53 copyright and training-data transparency obligations
Separately from the DSM Directive, the EU AI Act adds obligations specifically for providers of general-purpose AI (GPAI) models under Article 53(1)(c)–(d): a policy to comply with EU copyright law — including identifying and honoring Article 4(3) DSM rights reservations — and a publicly available, “sufficiently detailed summary” of the content used to train the model, using a template the European Commission published on 24 July 2025. These obligations apply from 2 August 2025 for models placed on the market after that date, and from 2 August 2027 for models already on the market before then; the AI Office’s full enforcement powers (recalls, mandated mitigations, fines) apply from 2 August 2026. Note that the EU’s 2026 “Digital Omnibus” process is actively renegotiating parts of the AI Act’s implementation — a provisional political agreement was reached in May 2026 — so these dates and mechanics are worth re-checking rather than treated as permanently fixed. Source: Article 53, EU Artificial Intelligence Act.
United Kingdom: CDPA 1988 s.29A
UK law’s TDM exception, s.29A of the Copyright, Designs and Patents Act 1988, is narrower than either EU exception: it permits copying for computational analysis only where the purpose is non-commercial research, the copier has lawful access, and a sufficient acknowledgement accompanies the copy. Like DSM Article 3, it cannot be excluded by contract. Commercial TDM — including most AI model training — has no UK statutory exception at all. The UK government ran a public consultation (December 2024–February 2025) that had proposed a broad opt-out-based commercial TDM exception; its 18 March 2026 report and impact assessment abandoned that proposal, stating there is currently insufficient evidence to justify changing the law, and maintaining the status quo pending further evidence. Sources: CDPA 1988 s.29A; UK government Report on Copyright and Artificial Intelligence (March 2026).
United States: case-by-case fair use, no statutory TDM exception
US copyright law has no dedicated TDM exception. Whether training on copyrighted material is lawful turns on the fact-specific, four-factor fair use test (17 U.S.C. §107), and 2025’s first rulings on AI training show it does not resolve uniformly:
- Bartz v. Anthropic (N.D. Cal., June 2025): the court held that training an LLM on lawfully acquired books was “exceedingly transformative” fair use — but that using pirated copies of the same books was not fair use, regardless of the transformative end use. Anthropic subsequently agreed to a settlement (roughly 482,000 works) that received preliminary approval, with a final fairness hearing set for 23 April 2026.
- Kadrey v. Meta (N.D. Cal., June 2025): a separate federal court reached a broadly similar fair-use conclusion on book-training claims around the same time.
- Thomson Reuters v. Ross Intelligence (D. Del., February 2025): the first US ruling to reject a fair-use defense for AI training data — but on a non-generative AI system that used extracted Westlaw headnotes to build a directly competing legal-research product. The court found the use not transformative and weighed market harm heavily against the defendant. The ruling was explicitly limited to non-generative AI and does not resolve how courts will treat generative model training generally.
The practical takeaway for a US-based research project: there is no bright-line statutory permission, and the case law so far suggests the outcome depends heavily on (a) whether the underlying copies were lawfully acquired and (b) how transformative the resulting use is — not on training-for-research purposes being automatically exempt.
Quick comparison
| Jurisdiction | Basis | Covers research TDM? | Covers commercial AI training? | Can rightsholders opt out / override? |
|---|---|---|---|---|
| EU | DSM Directive Art. 3 (research) / Art. 4 (general) | Yes, mandatory (Art. 3) | Yes (Art. 4), subject to opt-out | No override on Art. 3; yes on Art. 4 via rights reservation |
| UK | CDPA 1988 s.29A | Yes, non-commercial research only | No statutory exception | No override (contract terms restricting it are unenforceable) |
| US | Fair use, 17 U.S.C. §107 (case law only) | Case-by-case; no dedicated exception | Case-by-case; outcomes have diverged (2025 rulings) | Not applicable — no statutory exception to opt out of |
Documenting provenance in practice: what to put in the DMP
Regardless of which jurisdiction’s exception a project relies on, funders, publishers, and downstream reusers increasingly expect the same underlying documentation the EU AI Act now mandates for GPAI providers (a description of training content and its rights basis) even for research-only uses that aren’t legally required to publish one. A practical DMP entry or accompanying data statement for an AI/ML training dataset should cover:
- Sources — where each component of the corpus came from (named repositories, licensed corpora, web-scraped material, institutionally-held data), consistent with how a conventional dataset’s origin is already recorded via DataCite metadata or repository-level provenance fields.
- Rights basis per source — licence terms, a stated TDM exception relied on (and which one), or documented rightsholder permission. See licence and reuse license for how CASRAI’s Dictionary already frames this for conventional research data.
- Processing and filtering steps — deduplication, exclusion criteria, any subsetting that changes what’s actually represented in the trained model, aligned with the composition tracking described at Training data composition.
- Model/weights licensing — a resulting model’s own reuse terms, distinct from the training data’s terms (see Model weight licence).
- Known contamination or leakage risks — whether evaluation or held-out data may overlap with training sources (see Data leakage (training)).
Capturing this as structured, machine-actionable metadata in the DMP — rather than prose written after the fact — is the same argument CASRAI’s guidance already makes for machine-actionable DMPs generally (see machine-actionable DMP and the NIH vs. NSF Data Management Plans comparison for how funder-specific DMP formats already accommodate structured fields).
Frequently asked questions
Does a university’s own copy of licensed journal content qualify for TDM under a research exception?
Only where the institution has lawful access to the material in the first place — a licence that includes or does not prohibit TDM, or a statutory exception covering it (EU Art. 3, UK s.29A). “Lawful access” is a threshold condition in both the EU and UK exceptions, not an afterthought — check the underlying subscription or licence terms, not just copyright law in isolation.
Is there a US equivalent to the EU’s research-specific TDM exception?
No. US law relies entirely on fair use case law rather than a dedicated statutory exception, and — per the 2025 rulings above — a “for research” framing is not on its own a guarantee of fair use; the analysis is fact-specific to the use and the source of the copies.
If a project’s AI training corpus is entirely CC0 or CC-BY data, does any of this still apply?
The TDM-exception question becomes largely moot (permissively licensed data doesn’t need a copyright exception to be mined), but the provenance-documentation discipline still applies in full — reviewers, funders, and downstream reusers still need to know what the model was trained on and under what terms, independent of whether a legal exception was needed to get there.
Does documenting training data provenance satisfy the EU AI Act’s training-data summary requirement?
Only if it’s structured to the Commission’s published template and the project or institution is itself acting as a GPAI provider under Article 53 — most individual research projects aren’t. But building the habit of structured provenance documentation now means that obligation, if it ever applies, is a formatting exercise rather than a reconstruction project.
The litigation backdrop to all of this keeps moving. In July 2026 Hachette Book Group, Cengage, Elsevier and novelist Scott Turow sued Google over the use of books and scholarly content to train Gemini — one of several US cases testing whether training on copyrighted works is fair use, and a reminder that a model’s training corpus is a documentable, litigable fact.
Related CASRAI resources
- Research Data Management — cluster overview
- Dictionary: Training data provenance
- Dictionary: AI provenance
- Guide: Creative Commons Licenses for Research Data
- Guide: How to Choose an Open Data Repository
- Guide: NIH vs. NSF Data Management Plans
- Guide: How to Write a Data Availability Statement for Reproducibility







