Author hopes to throw the book at OpenAI, Microsoft with copyright class action
- Reference: 1700681354
- News link: https://www.theregister.co.uk/2023/11/22/openai_microsoft_copyright_complaint/
- Source link:
The crux of the case is all too familiar. An author is not happy that their work has been slurped as training data into OpenAI's text-generating models behind services such as ChatGPT. In this case, it is Julian Sancton, author of Madhouse at the End of the Earth, which documents an Antarctic polar expedition by a Norwegian steamship at the end of the 19th century.
Sam Altman set to rejoin OpenAI as CEO – seemingly with Microsoft's blessing [1]CONTEXT
According to the complaint, Sancton spent five years and tens of thousands of dollars on the book, secure in the knowledge that the US Copyright Act gives "exclusive rights" as well as "the rights to reproduce the copyrighted work[s]."
Sancton's complaint states: "This case is about defendants OpenAI and Microsoft's complete disregard for those exclusive rights.
"Defendants [OpenAI and Microsoft] have made commercial reproductions of millions, maybe billions, of copyrighted works without any compensation to authors, without a license, and without permission.
[2]
"In doing so, they have infringed on the exclusive rights of Plaintiff Sancton and other writers and rightsholders whose work has been copied and appropriated to train their artificial intelligence models."
[3]
[4]
In September, a [5]class action suit was launched by the Authors Guild, alleging that OpenAI's chatbots had been trained on work from such luminaries as George R R Martin and John Grisham. Another [6]lawsuit was brought in July 2023 by comedian Sarah Silverman.
[7]No more staff budget for UK civil service, but worry not – here's an incubator for AI
[8]Ex-OpenAI staff launch new chatbot – yup, it's Anthropic with Claude 2.1
[9]Microsoft dials back Bing after users manage to recreate Disney logo in fake AI-generated images
[10]Fake views for the win: Text-to-image models learn more efficiently with made-up data
Despite being filed on November 21, some elements of the complaint are out of date. While AI is a fast-moving field, that is nothing compared to the spin of the OpenAI revolving door. The complaint states: "The OpenAI-Microsoft relationship is so close, in fact, that OpenAI's former CEO Sam Altman and former Chief Scientist Greg Brockman just left the company to lead a new artificial intelligence research team at Microsoft."
As of today, [11]Altman is set to return to OpenAI. However, by tomorrow, this may have changed.
This latest lawsuit – alongside others alleging copyright theft and [12]privacy violations – is a reminder that the biz and its major investors, like other AI vehicles, face questions regarding where all the training data used in its models has come from.
[13]
In a statement, Justin A Nelson, partner with Susman Godfrey, lead counsel for Sancton and the proposed class, [14]said : "The commercial success of the ChatGPT products for OpenAI and Microsoft comes at the expense of non-fiction authors who haven't seen a penny from either defendant, much less a request for permission to use their works.
"This lawsuit seeks to hold OpenAI and Microsoft accountable for their refusal to pay nonfiction authors, and to prevent the companies from infringing on works in the future."
We have asked OpenAI and Microsoft to comment. ®
Get our [15]Tech Resources
[1] https://www.theregister.com/2023/11/22/sam_altman_openai_return/
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZV6IGeP41bSquFSuPZ2ytQAAABA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZV6IGeP41bSquFSuPZ2ytQAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZV6IGeP41bSquFSuPZ2ytQAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://www.theregister.com/2023/09/21/authors_guild_openai_lawsuit/
[6] https://www.theregister.com/2023/07/10/in_brief_ai/
[7] https://www.theregister.com/2023/11/22/uk_ai_incubator/
[8] https://www.theregister.com/2023/11/21/anthropic_claude_chatbot/
[9] https://www.theregister.com/2023/11/20/ai-in-brief/
[10] https://www.theregister.com/2023/11/22/texttoimage_models_mit/
[11] https://www.theregister.com/2023/11/22/sam_altman_openai_return/
[12] https://www.theregister.com/2023/06/28/microsoft_openai_sued_privacy/
[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZV6IGeP41bSquFSuPZ2ytQAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[14] https://www.susmangodfrey.com/news/lawsuit-seeks-to-hold-openai-and-microsoft-liable-for-rampant-theft-of-authors-works/
[15] https://whitepapers.theregister.com/
Re: So what about all the students reading books to write papers?
The argument is that reading a book as a human and processing the book in a process called "training" aren't the same thing. Thus, just because one is acceptable doesn't mean the other is. It gets philosophical when we start to ask what the model is really doing with the text it is ingesting, but it should be clear that I can't use whatever copyrighted information I want just by calling whatever my program is doing training.
Re: So what about all the students reading books to write papers?
> ...but it should be clear that I can't use whatever copyrighted information I want just by calling whatever my program is doing training.
It seems to me that the authors would have a point if their works were being duplicated by these systems and if a ChatBot did spew out verbatim passages from their works, then they would have a very real case.
Someone with a very good memory could read the works and recite them back into printed form and that (with fair use caveats) could constitute copyright infringement.
But unless it actually does, I don't really see a fundamental distinction between what people do when reading the book and what machines do through training, if these large language models are not merely storing coherent copies of the works.
As these models become more and more sophisticated and it increasingly seems like the difference between what people and machines do is narrowing, then that distinction is going to become harder and harder to make.
Re: So what about all the students reading books to write papers?
Not the same thing since these 'AIs' are not actually intelligent in the human sense at all. They simply use inductive logic to predict which words they could, or should use given conditions X and Y. To phrase it differently, they do not understand what these words 'mean' since they do not share in the human experience. A student may attempt to actually understand a given text, then contemplate its implications and integrate them into their own analysis of a situation or problem.
Zzzzzzzzzz
Sancton spent five years and tens of thousands of dollars on the book, secure in the knowledge that the US Copyright Act gives "exclusive rights" as well as "the rights to reproduce the copyrighted work[s]."
Non-starter. If I read Sancton's book and then set myself up as an expert on the expedition, giving lectures, taking money to help new expeditions learn from the expedition, proof reading other books to correct the spelling of the expedition leaders name, or in any other way make money from the knowledge I gained from reading Sancton's book, I don't owe him a penny.
And neither does OpenAI.
Look... I don't like the current pretend-AI (but really LLM) toys any more than the next guy. But for goodness sake let's hit these copyright claims trying to grab some of the AI-hype-money on the head. Facts can't be copyrighted. The information in Sancton's book is not copyrighted. In some cases the wording he uses may be, but it would be very hard to make a copyright claim on the wording of a fact unless the wording was very, very unusual.
Stop trying to jump on the bandwagon. Instead, take out your frustration by helping stop the hype - show how the book is much better than the so-called-AI.
So what about all the students reading books to write papers?
So what about all the students who read these books and write papers about the the information inside them? Is that stealing too? Just training (or reading as us wet brains call it) is what you do with books and media. It's how you learn and build upon other's works. Even citing small passages is enshrined in the Copyright Act (of the US, no clue on UK laws) as fair use.
But mostly I feel for all the monks who are out of jobs now that authors can just print their documents on a printer, without any human help! Woe!!!