Book Publishers Sue Google For Copyright Infringement Over Gemini AI Training (theguardian.com)
- Reference: 0184472746
- News link: https://yro.slashdot.org/story/26/07/15/2113245/book-publishers-sue-google-for-copyright-infringement-over-gemini-ai-training
- Source link: https://www.theguardian.com/books/2026/jul/14/publishers-sue-google-gemini-ai-training
> The publishers argue that Google repurposed books that had been supplied for limited services such as Google Books, Google Play Books and Google Scholar. Those services allowed Google to use the works in specific ways -- for example, to display searchable snippets or sell ebooks -- but not, the lawsuit claims, to copy them for training commercial AI products. "Desperate to maintain its online dominance, Google abandoned its early motto of 'Don't be evil' and engaged in one of the most prolific infringements of copyrighted materials in history," the suit [2]states (PDF).
>
> According to the complaint, the tech company made copies of copyrighted books to train Gemini without permission or payment, despite internal discussions acknowledging the legal risks. The filing claims Google flagged internally that it could face "$10Bs-$100Bs in potential fines" for using texts provided by publishers for Google Play Books. The publishers say Google's actions are harming authors and the wider publishing industry, arguing that AI-generated content could negatively impact book sales.
>
> It notes that, for example, Gemini could generate "a 100-page murder mystery set in a quiet seaside town filled with secrets, that substitutes for an original copyrighted murder mystery on which Gemini trained" in 20 minutes for 39 cents. "No publisher or author can compete with that." The lawsuit names a number of specific books that the publishers allege were among the copyrighted works used without permission, including NK Jemisin's The Fifth Season, and Lemony Snicket's Who Could That Be at This Hour?
[1] https://www.theguardian.com/books/2026/jul/14/publishers-sue-google-gemini-ai-training
[2] https://publishers.org/wp-content/uploads/2026/07/Hachette-v.-Google-Dkt.-1-Complaint2.pdf
Dictionaries Mysteriously Not Sued (Score:1)
Why is it, that dictionary publishers are not sued? Lexicographers employ much the same methods in studying/analysis of words of copyrighted works, but rather, often "by hand," to develop citations to their corpus (body of evidence from which dictionary statistics are drawn in order to formulate a dictionary entry), and AI seems to even be modeled after the way lexicography has done this very thing for eons, but yet dictionary publishers never seem to be sued about it.
Re: Dictionaries Mysteriously Not Sued (Score:2)
Because the dictionary publishers didn't copy all those works into their training corpus
Re:Dictionaries Mysteriously Not Sued (Score:4, Informative)
Dictionary publishers do get sued. In 2001, the "New Oxford American Dictionary" added ghost (fake) word "esquivalience" to their boko and sued several online dictionaries that included the word (who only could have added through bulk copying the first dictionary).
In 1998, Larousse and Robert (two well-known dictionary publishers) sued Maxidico for plagiarism due due to striking similarities in definitions, including same mistakes/typos. Maxidico was sentenced to the equivalent of 1.5 million euros in damages (of the money of the time) and filed for bankrupcy.
Re: (Score:3)
Dictionary publishers have never been accused of downloading massive torrents of pirated copies of books and processing them.
Google on the other hand HAS been accused of that, and the decade of litigation related to that ultimately rules that the limited things google was doing with it was fair use. The dictionary companies are likely paying for enhanced access to that google data now.
The AI companies are singing the same fair use tune, but its really quite different. Google was doing it (at the time) to al
Re: (Score:2)
No, its not. In fact, I'd wager that they only have one copy of each unique word in their database*.
* - you have to assume that they also have development, testing, and pre-production copies of that same database.
Re: (Score:2)
Who is they? And what database are you referring to?
Dictionaries often contain quotes from original source material, or references to where it was first used, or first used a certain way.
Re: (Score:1)
No. It is not copyright infringement, and there's no reason to hold copyright so sacred anyway. Are you seriously wanting to protect hundred year old fairy tails from being retold? Information wants to be free.
Re: (Score:2)
"No. It is not copyright infringement"
Go ahead, prompt for that story and publish your own 'moonlit princess". It is not a court case you'd win; the details taken from the Disney version are beyond excessive.
" and there's no reason to hold copyright so sacred anyway. Are you seriously wanting to protect hundred year old fairy tails from being retold?"
That's an entirely separate discussion. Legally it is infringement. Whether it should be is completely separate question, or how long it should be are separate
Training is fair use, but... (Score:2)
Training is "inherently transformative", and thus protected as fair use. BUT... Google should lose this.
The publishers have a contract with Google that spells out the specific purposes for which the materials are to be used. These are not books that Google purchased off-the-shelf. They were provided by the publishers for that specific use. Any other use - even an otherwise legal use is a violation of that contract.
Even if Google argues fair use in training their AI system, they violated the contract. Pay
Re: (Score:1)
> Training is "inherently transformative", and thus protected as fair use. BUT... Google should lose this.
> The publishers have a contract with Google that spells out the specific purposes for which the materials are to be used. These are not books that Google purchased off-the-shelf. They were provided by the publishers for that specific use. Any other use - even an otherwise legal use is a violation of that contract.
> Even if Google argues fair use in training their AI system, they violated the contract. Pay up.
There are [1]four criteria that are used to help determine whether something is fair use [cornell.edu] (in the US) and Google seems to fall on the wrong side of three on them:
(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
(2) the nature of the copyrighted work;
(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
(4) the effect of the use upon the potential market for or value of th
[1] https://www.law.cornell.edu/uscode/text/17/107
Re: (Score:2)
> Training is "inherently transformative", and thus protected as fair use.
You should stop posting on the law until you understand it better. Go watch some lawyer videos on fair use or something. The Wikipedia article is good. Your concept is wrong.
You cannot copyright a genre (Score:1, Informative)
You can copyright specific series of phrases or words in a specific order, but you cannot copyright/trademark a style of writing; this is why there are so many Call of Duty style games, otherwise the makers of the early video game works like Wolfenstein 3D (or earlier) could sue the pants of every first-person-shooter who resembles it. You copyright the specific art, the specific wording, the performed audio, potentially elements of the code, but almost all such games formulaically have menus where you choo
These publishers are part of the problem (Score:2, Offtopic)
Elsevier - the company paywalling research journals and charging $50 bucks to see one article? A post-scarcity future has no copyrights. Did they ever mention copyrights on Star Trek???
These LLM models have taken ALL of humanities knowledge and ALL of humanity should benefit, as seen previously on Slashdot: Bernie Sanders Unveils $7 Trillion Plan To Give Americans Control of AI Industry [1]https://yro.slashdot.org/story... [slashdot.org]
[1] https://yro.slashdot.org/story/26/06/18/1914206/bernie-sanders-unveils-7-trillion-plan-to-give-americans-control-of-ai-industry
Where's the payout for coders? (Score:3)
Seriously, why should book authors get more compensation for their work being trained on than the rest of us?
Re: (Score:2)
Nobody's going to lift a finger on our behalf. That's what the thieves are counting on. Which I guess explains the vigorousness of the music industry's response to "pirating" back in the day.
Re: (Score:2)
Well, I would like to be compensated for the times my work got imported into AI, too.
Re: (Score:1)
I taught Gemini how nose-picking works. They owe me!
Re: (Score:2)
book authors never agreed to make their works freely available to anyone, unlike open source coders.
So, look and feel redux? (Score:1)
Google AI:
It really is 'look and feel' redux, but shifted from the graphical interface to the cognitive interface. In the 90s, the courts realized you couldn't lock up the abstract 'feel' of a desktop environment because doing so would stifle software evolution.
Now, publishers are trying to protect the traditional boundaries of book text. But Gemini isn't copying pages; it is learning the 'look and feel' of well-curated facts and analytical structure to create a completely new kind of dynamic interface. It'
Typically american⦠(Score:2)
Typical American corporation, I can take whatever I want, if youâ(TM)re not happy, sue me!
Consensus cracking (Score:1)
There's no reason Slashdot would suddenly come down hard on "infringement" by AIs except misguided general anti-AI sentiment and consensus cracking. Some people don't like AI and have seized on this as a way to stop it. Some other people are paid to say these things, to create the illusion that there is consensus for strong copyright laws. But strong copyright laws hurt almost everyone and have gone far, far beyond their original justifications. Rent seeking owners of publishers are holding our culture host
Not sure what the answer is? (Score:2)
You have to admit the Book Publishing business looks pretty buggy whippy atm.
And related to Authors and others, yea they got robbed, but when it comes to LLM generated material not sure how it gets stopped now.
You are already seeing huge mountains of garbage burying things of value.
And going forward anything put on the Internet will be scraped if it has a hint of value.
Re: (Score:2)
You can't stop the LLM if it's published... but you can sue the company that scraped data it was not legally entitled to scrape, and the legal sanctions should involve destruction of the collected archive of training data and all copies of the resulting LLM as well as a financial penalty that is sufficiently large to dissuade future repetitions of the offense.
"But that would harm our bottom line" is not an acceptable defense against this. Don't steal. It's easy.
Of course, stealing's easier if you're a meg
Judgment the world wants (Score:2)
The judgment the world wants:
Google; You wre evil and will be broken up.
Copyright Cartel: 95-year copyright is an abomination. That goes down to a reasonable five years, and one year for journals.
Good (Score:2)
They like any of the other should pay for what they have taken. In an ideal world they'd be forced to remove it from their data until they have suitable permission but I don't foresee that happening or if it's even possible without retraining the entire model.