If you're going to train AI on our books, at least pay us, authors tell Big Tech
- Reference: 1689701234
- News link: https://www.theregister.co.uk/2023/07/18/ai_in_brief/
- Source link:
Large language models are trained on large amounts of text scraped from the internet. Hundreds of thousands of books hosted on websites have been ingested without writers' permission. Now many of those writers are speaking out against having their work ripped off by computers.
"Generative AI technologies built on large language models owe their existence to our writings," begins the [1]letter addressed to the CEOs of OpenAI, Alphabet, Stability AI, Meta, IBM, and Microsoft. "These technologies mimic and regurgitate our language, stories, style, and ideas. Millions of copyrighted books, articles, essays, and poetry provide the 'food' for AI systems, endless meals for which there has been no bill."
[2]
"You're spending billions of dollars to develop AI technology. It is only fair that you compensate us for using our writings, without which AI would be banal and extremely limited."
[3]
[4]
Mary Rasenberger, CEO of the Authors Guild, [5]told NPR that the letter was written to try and get the companies to settle with writers without having to take matters into court. "Lawsuits are a tremendous amount of money … They take a really long time," she said. Other writers, however, have been more aggressive and have [6]sued those they see as having stolen their work.
LLaMa for profit
Meta will reportedly release a new version of its large language model, LLaMA, that supports commercial use in an attempt to compete with rival AI developers.
The social media giant often releases its models for academic research, and was criticized for not being as open as it claimed by preventing developers using LLaMA for commercial applications.
Zuckerberg's biz is reportedly looking at how it might be able to charge enterprises to fine-tune the model on its own custom data – but may not end up charging users at all, [7]according to the Financial Times .
[8]
The hope is that by sharing its models with more developers, Meta will be able to take some of the shine away from competitors OpenAI, Google, and Microsoft. The rumor was [9]first reported by The Information last month.
[10]Sarah Silverman, novelists sue OpenAI for scraping their books to train ChatGPT
[11]Google, DeepMind accused of 'stealing the internet' to create Bard AI chatbot
[12]OpenAI's ChatGPT may face a copyright quagmire after 'memorizing' these books
OpenAI inks data deals with Associated Press and Shutterstock
Facing an onslaught of copyright lawsuits, OpenAI has announced partnerships with the Associated Press and Shutterstock to license their content for training AI models. although that does little for the thousands of book authors (to say nothing of the millions of authors of other written works).
In this more "enlightened" scenario, OpenAI will get its hands on an archive of text dating back to 1985 published by the non-profit news agency, and AP will get access to the startup's "technology and product expertise" in return.
Last week in a statement, Kristin Heitmann, AP senior vice president and chief revenue officer, [13]said : "We are pleased that OpenAI recognizes that fact-based, nonpartisan news content is essential to this evolving technology, and that they respect the value of our intellectual property."
Both entities will team up to look for "potential use cases for generative AI in news products and services." The deal follows a similar [14]announcement from stock image vendor Shutterstock the week before that revealed OpenAI had signed a six-year agreement to license its content.
[15]
OpenAI will use the data to train its generative AI systems, while Shutterstock can continue using its technology to power tools like its AI Image Generator. The license agreement allows OpenAI to collect and obtain data with permission, while they are compensated for sharing their resources. The financial details of both deals were not disclosed. ®
Get our [16]Tech Resources
[1] https://actionnetwork.org/petitions/authors-guild-open-letter-to-generative-ai-leaders
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZLcLhNjLf8oFKjHHFkL8dgAAAVE&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZLcLhNjLf8oFKjHHFkL8dgAAAVE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZLcLhNjLf8oFKjHHFkL8dgAAAVE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://www.npr.org/2023/07/17/1187523435/thousands-of-authors-urge-ai-companies-to-stop-using-work-without-permission
[6] https://www.theregister.com/2023/07/10/in_brief_ai/
[7] https://www.ft.com/content/01fd640e-0c6b-4542-b82b-20afb203f271
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZLcLhNjLf8oFKjHHFkL8dgAAAVE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://www.theinformation.com/articles/meta-wants-companies-to-make-money-off-its-open-source-ai-in-challenge-to-google
[10] https://www.theregister.com/2023/07/10/in_brief_ai/
[11] https://www.theregister.com/2023/07/12/google_alphabet_deepmind_bard_complaint/
[12] https://www.theregister.com/2023/05/03/openai_chatgpt_copyright/
[13] https://apnews.com/article/openai-chatgpt-associated-press-ap-f86f84c5bcc2f3b98074b38521f5f75a
[14] https://www.prnewswire.com/news-releases/shutterstock-expands-partnership-with-openai-signs-new-six-year-agreement-to-provide-high-quality-training-data-301873298.html
[15] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZLcLhNjLf8oFKjHHFkL8dgAAAVE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[16] https://whitepapers.theregister.com/
Re: There's another risk: repurposing copyrighted data
Things will get very interesting when someone uses a LLM to create the lyrics to a song, the record companies will ensure this goes to court…
"It is only fair that you compensate us for using our writings, without which AI would be banal and extremely limited."
So all of the authors are top rank, none of the 90% who, following Sturgeon's Law, would only make the AI even more banal?
(Yes, saw there ARE well-known names: especially near the start of the list)
If they don't think the authors' work has any benefit, they are free to leave it out of the training data. That they have not suggests they think there is value in having that text in there, and they are using that value to make money. It's not on us to decide how much value they are getting from any given book, but on them to decide whether they are willing to pay for the use of copyrighted data they don't own. They can decide to exclude something because it is not available for sale, because it isn't worth as much as is being asked, or because they think it will be detrimental.
Never mind the quality, feel the width
> If they don't think the authors' work has any benefit, they are free to leave it out of the training data
True - except there doesn't seem to be any indication that they are spending all of the money/effort/time required to evaluate the content of their entire training set: after all, doing so would only reduce the amount of bulk material used and part of the boasting is how much text was used[1]!
Although, having said that, you can also feed the training with a good dose of negatives: "please don't write stuff like this". So even after categorising every bit of their dataset they'll still use it all, just maybe not in the way that the author's would like (they aren't happy now, but if they find that they are on the Naughty Training List they're likely to get really upset!)
[1] some say 5GB for GPT-1, 40GB for GPT-2, 600GB for GPT-3
They only want to whip the LLaMa's ass
> The social media giant often releases its models for academic research, and was criticized for not being as open as it claimed by preventing developers using LLaMA for commercial applications.
Huh? Open generally just starts with "you can read it", verify it and so on - *some* licences (including the most famous ones, at least to the audience here) then go on to say that you can also *use* it, with or without other restrictions. Such as, not for commercial use.
This isn't complaining about lack of openness, it is just complaining about not being allowed to exploit.
Which was probably a Good Thing - these LLMs are already being shoved into use in places where they are not fit for purpose[1], no need to encourage more of that.
Fingers crossed, the newer LLaMa, if it is released for commercial use, will at least have benefitted from being poked by academics and will, compared to its predecessors, have less chance of making a total mockery out of everyone deploying it.[2]
[1] and, yes, you can indeed argue that applies to every use made of them so far; I'll agree for the well-hyped uses by the Big Players, certainly.
[2] or Meta have realised they'll never get any better, so soak everyone in one fell swoop before the word gets out; place your bets now.
There's another risk: repurposing copyrighted data
Given that these models are trained on data and texts that are under copyright (and patents, and other legal protections), any output will have to be checked for required contribution distribution.
But I think it may get worse.
Once that output is used by other parties, that content again enters under the various legal umbrellas that exist - and the only winners in the fight to then untangle the mess are lawyers.
Not good.
Anyone publishing information should really start to check the legal protections they have - if you check Google's conditions, or Facebook's, you will find that you signed away all your rights. They've been working on this for a long time - you're not just milked for personal details, but also for free contents. Now them chickens are coming home to roost..