Authors Guild sues OpenAI for using Game of Thrones and other novels to train ChatGPT
- Reference: 1695281468
- News link: https://www.theregister.co.uk/2023/09/21/authors_guild_openai_lawsuit/
- Source link:
Named plaintiffs in the copyright infringement class action lawsuit – filed in the Southern District of New York for copyright – include David Baldacci, Mary Bly, Michael Connelly, Sylvia Day, Jonathan Franzen, John Grisham, Elin Hilderbrand, Christina Baker Kline, Maya Shanbhag Lang, Victor LaValle, George R.R. Martin, Jodi Picoult, Douglas Preston, Roxana Robinson, George Saunders, Scott Turow, and Rachel Vail.
The [1]complaint [PDF] argues that OpenAI's services "endanger fiction writers' ability to make a living, in that the large language models allow anyone to generate – automatically and freely (or very cheaply) – texts that they would otherwise pay writers to create."
[2]
The complaint points out that ChatGPT has successfully been prompted to create a "detailed outline for a prequel book to A Game of Thrones … using the same characters from Martin's existing books in the series A Song of Ice and Fire ." Similar results were possible for the other authors who have joined the suit.
[3]
[4]
ChatGPT's ability to do so is problematic, given that the authors did not authorize OpenAI to access their works.
The complaint states that OpenAI has admitted to using datasets named "Books1" and "Books2" to train its large language models, but hasn't disclosed their content. The plaintiffs suspect pirate books have made their way into OpenAI training data.
[5]
"The growth in power and sophistication from GPT-3 to GPT-4 suggests a correlative growth in the size of the 'training' datasets, raising the inference that one or more very large sources of pirated ebooks discussed above must have been used to 'train' GPT-4," the complaint argues, adding "There is no other way OpenAI could have obtained the volume of books required to 'train' a powerful LLM like GPT-4."
Actually, the complaint does mention one other way: paying for the content used to train ChatGPT. But the suit alleges OpenAI never thought to do so, and quotes CEO Sam Altman's testimony to Congress that he believes in Copyright and has paid for some training data.
[6]Textbook publishers sue shadow library LibGen for copyright infringement
[7]Don't worry, folks. Big Tech pinky swears it'll build safe, trustworthy generative AI
[8]OpenAI boss Sam Altman receives Indonesia's first golden visa
[9]OpenAI urges court to throw out authors' claims in AI copyright battle
"For fiction writers, OpenAI's unauthorized use of their work is identity theft on a grand scale," stated Authors Guild CEO Mary Rasenberger.
"Fiction authors create entirely new worlds from their imaginations – they create the places, the people, and the events in their stories," she added, before lamenting "People are already distributing content generated by versions of GPT that mimic or use original authors' characters and stories. Companies are selling prompts that allow you to 'enter the world' of an author's books. These are clear infringements upon the intellectual property rights of the original creators."
The plaintiffs want "damages for the lost opportunity to license their works, and for the market usurpation Defendants [OpenAI] have enabled by making Plaintiffs unwilling accomplices in their own replacement; and a permanent injunction to prevent these harms from recurring."
[10]
The Register has asked OpenAI for comment and will update this story if we receive a substantial reply. ®
Get our [11]Tech Resources
[1] https://regmedia.co.uk/2023/09/21/authors_guild_openai_complaint.pdf
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZQwUQuA9UKt1AOsBa9CVYAAAAJc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZQwUQuA9UKt1AOsBa9CVYAAAAJc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZQwUQuA9UKt1AOsBa9CVYAAAAJc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZQwUQuA9UKt1AOsBa9CVYAAAAJc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2023/09/18/science_publishers_sue_libgen/
[7] https://www.theregister.com/2023/09/12/nvidia_adobe_palantir_ai_safety/
[8] https://www.theregister.com/2023/09/12/sam_altman_indonesia_gold_visa/
[9] https://www.theregister.com/2023/08/31/openai_class_action_fair_use/
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZQwUQuA9UKt1AOsBa9CVYAAAAJc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[11] https://whitepapers.theregister.com/
Looks like they are going after "the reader" not "the poster" of copyright material ...
Much more money in litigation against multiple readers (or one rich one) instead of shutting down the thief.
Reading the published material is not a crime, using the material for any business purpose or publishing similar works to the detriment of the Author is.
It's the same law Disney use when not getting paid by someone selling micky mouse wallpaper.
but the AI read the works, it's not "using the works"
Only the living can sue.
Which leaves OpenAI safe to rip off all the dead authors.
Can it write a non derivative Culture or Discworld novel with new characters to the same level as Banks & Pratchett?
Even if it could I wouldn't pay more than printing cost anyway.
*Both cruelly taken far too early
Re: Only the living can sue.
The dead can't sue, but their publishers can. Being dead doesn't put your works in the public domain; not immediately, at least.
All authors started as readers
There are two aspects on the copyright attack against AI.
The first is that the models are trained on existing texts. This training involves copying, and is therefore "forbidden". The same argument can be made against every author there ever was. They all learned the trade by reading texts of other authors. Copyright law works by preventing the use of protected works in publishing. The benchmark is whether the copied work is identifiable in the new work. That is most certainly not the case here. The fact that chatGPT can write a protected work is not different from MS Word being able to write a protected work. The courts have already decided that AI cannot produce works on its own. And I am certainly in my rights to write, eg, fan-fiction for myself. Unless I publish it, I can write whatever I like.
The second part is, like Andy the Hat writes, that the authors claim OpenAI used illegal copies for their training. As the authors seem to be unable to point out evidence of which pirated copies were used, this seems a little desperate.
I think the main point of the authors comes from this line:
> The complaint [PDF] argues that OpenAI's services "endanger fiction writers' ability to make a living, in that the large language models allow anyone to generate – automatically and freely (or very cheaply) – texts that they would otherwise pay writers to create."
The proverbial buggy whip manufacturers that want to stop Henry Ford destroying their revenues. If AI can write you a story as good as the authors can, why pay the authors? Indeed, why pay buggy whip manufacturers when you do not need a buggy anymore?
Even if the authors can make their argument stick and force AI to refrain from using books under copyright. That won't stop AI from writing books. The Iliad and Odyssey are some of the oldest surviving adventure novels and can be a very good start to write up everything from Game of Thrones to Space operas. And then we have not even started with Shakespeare. I am pretty sure AI can be nudged into combining the old texts with the new world and get us the books we want.
And that is before a user can feed a digital book into an AI and asks it to write a sequel.
"If AI can write you a story as good as the authors can, why pay the authors?"
The outrage here is that the machines are no longer coming for the jobs of the working class who toil and sweat and use their hands. Now they're coming for the comfortable middle class who've got (to quote Pratchett) an indoor job with no heavy lifting. And I think the aforementioned working class aren't going to be brimming with sympathy for the keyboard jockeys who see their livelihood going the way the coal mines went in the 80s.
They're not coming for the GOOD ones - not yet. So far the only ones they can actually replace are the derivative hacks... but most authors, even the good ones, start out as somewhat derivative until they find their voice. Pratchett's "Strata" was a transparent parody of Niven's "Ringworld", and clearly a sort of practice run at a Discworld. And even the biggest Pratchett fan will admit it's not as good as most of what followed (I happen to love it for what it is.)
But I think if someone were able to synthesise a new Culture novel (not a parody, not a reboot, an actual new Culture novel)... I think I'd want it. I'd dearly like IMB back, but if a LLM (with help, presumably, from someone with the right prompts) could make more work that is aesthetically equal to what already exists... why wouldn't you want it? Just out of principle?
> The outrage here is that the machines are no longer coming for the jobs of the working class who toil and sweat and use their hands.
Funnily enough, these jobs now look safer than a lot of so called "knowledge worker" jobs. Why pay for a photographer when you can have glamour shots from a short prompt and a crappy selfie?
I think this is a loosing fight. What today requires a data center will be in reach of individual users in few year's time. Then we will see self hosted LLMs.
"one or more very large sources of pirated ebooks"
I would have liked to by a fly on the wall of the meeting that decided to go get pirated material to use as training data.
Mgr - "Okay, guys, we have this ginormous potential waiting on training data. Where can we get that ? Ideas ?"
Mkting - "Well, we could strike deals with the Project Gutenberg website, they've got plenty of free books. I'm sure they'd be willing to help."
Mgr - "How much would that cost ?"
Mkting - "It's free for the customer, but we'd need a deal where we can get stuff in bulk. Shouldn't cost more than a couple thousand."
Mgr - "How long would that take ?"
Mkting - "I guess a month or two to negociate the deal and have a contract written up."
Mgr - "Too long. We need to move forward now . Any other ideas ?"
Dev - "Well, I know this site where we can get just about everything. All I'd need to do is write a script to automate the downloads."
Mgr - "What about the contract ?"
Dev - "Um, well, there isn't any. It's BitTorrent-like, you just go choose and it drops in."
Mgr - "And we can get recent stuff, no problem ?"
Dev - "Well yeah. Pirates love recent stuff."
Mgr - "Pirated ? So no contract and no money ?"
Dev - "Nope. And it's untraceable."
Mgr - "Go for it !"
A Song Of Ice and Fire
Maybe OpenAI could complete ASOIAF because it doesn't look like GRRM will.