Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Editorial · CASRAI · Compliance and regulatory

Publishers, Turow Sue Google Over Gemini Training

Hachette, Cengage, Elsevier, and Scott Turow sued Google in SDNY, alleging Gemini trained on Google Books, Play Books, and Scholar without authorization.

Published 23 Jul 2026· 6 minute read

Hachette Book Group, Cengage Learning, Elsevier, novelist Scott Turow, and his company S.C.R.I.B.E. have filed a copyright lawsuit against Google in the U.S. District Court for the Southern District of New York, first reported in the trade and legal press on July 14, 2026 (some outlets report the complaint was docketed July 10, with the plaintiffs’ public announcement following July 13). Multiple outlets identify the matter as Hachette Book Group, Inc. et al. v. Google LLC. This is a separate, later case from the copyright suit some of the same publishers filed against Meta in May 2026 (below) — the defendant, the alleged source of the training material, and several of the legal theories differ.

What the complaint alleges

The plaintiffs allege Google used works submitted to Google Books, made available through Google Play Books, and indexed via Google Scholar to train its generative-AI models, specifically Gemini, without the additional permission they say that use requires. The core theory reported across coverage: publishers and authors authorized Google, under earlier agreements and settlements tied to the Google Books project, to scan and display limited snippets of copyrighted works for search purposes — not to repurpose the full text of those works as training data for a generative-AI product.

The complaint reportedly covers an unusually broad range of material for a single suit: fiction, nonfiction, children’s books, memoirs, and poetry from the trade side (Hachette, Turow), educational textbooks (Cengage), and peer-reviewed scientific and scholarly journal articles (Elsevier) — the last category tied specifically to content indexed through Google Scholar. That scholarly-journal dimension is what puts this suit in research-administration and scholarly-publishing territory rather than purely a trade-publishing story.

Plaintiffs also allege Google removed or altered copyright management information (author, title, and rights notices) from works before using them for training, in order to obscure the origin of the material — a claim that, if it tracks the pattern of similar allegations in other AI-training suits, would invoke the Digital Millennium Copyright Act’s copyright-management-information provision (17 U.S.C. § 1202) alongside the underlying copyright-infringement claims.

What the plaintiffs are seeking

Reported relief sought includes statutory damages, an injunction against continued use of the works to train Gemini, and destruction of the allegedly unauthorized training copies. The suit is framed as a proposed class action on behalf of a broader group of rightsholders whose works were submitted to Google Books, Play Books, or indexed by Scholar.

How this relates to the May 2026 suit against Meta

Several of the same plaintiffs — Elsevier, Cengage, Hachette, and Scott Turow — already have a separate, earlier case pending: a class action filed May 5, 2026 against Meta and its CEO personally, joined there by Macmillan and McGraw Hill, alleging Meta trained its Llama models on more than 267 terabytes of copyrighted material obtained from shadow-library sites such as LibGen and Anna’s Archive. The two cases share several plaintiffs and a common underlying grievance — unauthorized use of copyrighted text to train a large language model — but rest on different factual theories: the Meta suit centers on allegedly pirated source material, while the Google suit centers on works Google obtained through legitimate, scope-limited arrangements (Google Books scanning, Play Books distribution, Scholar indexing) and is alleged to have used beyond that scope. Readers should treat them as two distinct pieces of litigation, not one case reported twice.

Both suits land in an AI-training copyright landscape that already has real precedent, none of it fully settled. Two federal court decisions in 2025 (in the Northern District of California, in litigation involving Anthropic and separately Meta) found that training an AI model on lawfully acquired copyrighted text can qualify as fair use, while treating the acquisition of pirated copies as a separate, unprotected question — a distinction plaintiffs’ counsel in the newer suits appear to be leaning on directly. Separately, Anthropic agreed to a reported $1.5 billion settlement in 2025 over its use of pirated books, described at the time as the largest publicly reported copyright settlement in the U.S. to date. None of this resolves how the Google suit’s specific facts — authorized-but-scope-limited access, rather than piracy — will be treated; that is likely to be the central legal question in Hachette v. Google itself.

Why this matters for research administration and scholarly publishing

The Google Scholar dimension of this suit is the part most directly relevant to CASRAI’s audience. If a court finds that indexing arrangements built for discovery and search do not implicitly license generative-AI training use, that has direct implications for how institutions, publishers, and repositories negotiate copyright and text-and-data-mining terms going forward — see CASRAI’s guide on AI training data provenance, copyright, and TDM exceptions for research for how that scope question is already playing out in licensing negotiations. It also sits alongside the wider shift toward explicit, negotiated AI-licensing deals between large publishers and AI developers — see CASRAI’s earlier coverage of scholarly publisher AI licensing deals in 2026 — which this litigation, win or lose, is likely to accelerate as the default way publishers try to control the terms rather than litigate them after the fact. Journal and manuscript-level AI policy is a related but distinct question; see CASRAI’s guide to journal and publisher policies on generative AI in manuscripts.

Frequently asked questions

Is this the same lawsuit as the publishers’ case against Meta?

No. The Meta case (filed May 5, 2026, Elsevier, Cengage, Hachette, Macmillan, McGraw Hill, and Scott Turow as plaintiffs) is a separate, earlier suit alleging Meta trained Llama on pirated copies obtained from shadow libraries. This Google case, filed in July 2026 in the Southern District of New York, alleges Google exceeded the scope of legitimate access it had to works through Google Books, Play Books, and Scholar. They involve overlapping plaintiffs but different defendants, different alleged source material, and different legal theories.

Does this suit involve academic journal content specifically?

Yes — reporting on the complaint specifically names peer-reviewed scientific and scholarly journal articles, tied to Elsevier as a plaintiff and to content indexed through Google Scholar, alongside the trade and educational-publishing material from the other plaintiffs.

Has a court ruled on the merits yet?

No. As of this writing the case is newly filed; no fair-use or infringement ruling has been reported. Prior 2025 rulings in separate AI-training cases (involving Anthropic and Meta) found training on lawfully acquired text can be fair use while treating pirated acquisition as a separate issue — a distinction likely to matter for how this differently-postured case proceeds, but not a ruling on this case itself.

This page tracks a fast-moving, newly filed piece of litigation. Details — case number, exact filing date, and claims — are based on multiple corroborating news reports at time of writing and may be revised as primary court filings and further reporting become available.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →