News: 1700051479

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

To pay or not to pay for AI's creative 'borrowing' – that is the question

(2023/11/15)


In the UK's Parliament this week, Microsoft and Meta ducked the question of whether creators should be paid when their copyrighted material is used to train large language models.

The tech titans, with combined revenues well in excess of $200 billion, were being grilled by the House of Lords Communications and Digital Committee when the copyright question came into focus.

In September, the Authors' Guild, a trade association for published writers, and 17 authors [1]filed a class-action lawsuit in the US over OpenAI's use of their material to create its LLM-based services.

[2]

OpenAI CEO Sam Altman has since said the company would cover its clients' legal costs for copyright infringement suits rather than remove the material from its training sets.

[3]

[4]

Microsoft has [5]invested $13 billion in OpenAI. It has an extended partnership with the machine learning developer, powering its workloads on the Azure cloud platform and using its models to run the Copilot automated assistant.

Speaking to the Lords yesterday, Owen Larter, director of public policy at Microsoft's Office of Responsible AI, said: "It's important to appreciate what a large language model is. It's a large model trained on text data, learning the associations between different ideas. It's not necessarily sucking anything up from underneath."

[6]

He said there should be a "framework" to provide some protection for copyrighted material and Microsoft would assume responsibility for any infringement by its LLM-based systems. But he also said Microsoft supports the recent [7]Valance report into "pro-innovation" AI law in the UK which advocates for text and data exceptions in training models.

But Donald Michael, Lord Foster of Bath, pressed Larter on whether he would accept that if a company uses copyrighted material to build an LLM for profit, the copyright owner should be reimbursed.

The Microsoft director said: "It's really important to understand that you need to train these large language models on large data sets if you're going to get them to perform effectively, if you're going to allow them to be safe and secure … There are also some competition issues [in making sure] that training of large models is available to everyone. If you go too far down a path where it's very hard to obtain data to train models, then all of a sudden, the ability to do so will only be the preserve of very large companies."

[8]

[9]Litigation is already under way to address how training data sets [10]Books1 , Books2, and Books3, which effectively pirate copyrighted material, have been used to help build popular LLMs.

[11]Google DeepMind's GraphCast AI weather predictor looks fascinating on paper but ...

[12]UnitedHealthcare's broken AI denied seniors' medical claims, lawsuit alleges

[13]YouTubers kindly asked to mark their deepfake vids as Fake Fakey McFake Fakes

[14]AI chemist creates catalysts to make oxygen using Martian meteorites

Meta is behind the [15]Llama 2 LLM , which scales up to 70 billion parameters. The social media giant has promoted the model as open source, although FOSS purists point to some caveats in its approach.

Speaking to the Lords, Rob Sherman, vice president and deputy chief privacy officer for policy at Meta, said the company would comply with the law.

But he added that "maintaining broad access to information on the internet and information including for the use in innovation like this is quite important. I do support giving rights holders the ability to manage how their information is used.

"I'm a little bit cautious about the idea of forcing companies that are building AI to enter into bespoke agreements with individual rights holders or an order to pay for content that doesn't have economic value for them."

Last week, Dan Conway, CEO of the UK's Publishers Association, told the committee that large language models were infringing copyrighted content on an "absolutely massive scale."

"We know this in the publishing industry because of the Books3 database which lists 120,000 pirated book titles, which we know have been ingested by large language models," he said. "We know that the content is being ingested on an absolutely massive scale by large language models. LLMs do infringe copyright at multiple parts of the process in terms of when they collect this information, how they store this information, and how they how they handle it. The copyright law is being broken on a massive scale."

At the same hearing, Dr Hayleigh Bosher, reader in intellectual property law at Brunel University London, said she did not represent tech firms or content creators and offered up a neutral's perspective.

"The principle of when you need a licence and when you don't is clear," she said, "and to make a reproduction of a copyright-protected work without permission would require a licence or would otherwise be infringement. That's what AI does at different steps of the process: The ingestion, the running of the program, and potentially even the output.

"Some AI and tech developers are arguing a different interpretation of the law. I don't represent either of those sides. I'm a copyright expert, and from my position, understanding of what copyright is supposed to achieve and how it achieves it, you would require a licence for that activity." ®

Get our [16]Tech Resources



[1] https://www.theregister.com/2023/09/21/authors_guild_openai_lawsuit/

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZVT5N091fb8cciDR9a-t8QAAAFI&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZVT5N091fb8cciDR9a-t8QAAAFI&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZVT5N091fb8cciDR9a-t8QAAAFI&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://www.theregister.com/2023/01/23/microsoft_openai_applications/

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZVT5N091fb8cciDR9a-t8QAAAFI&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://www.gov.uk/guidance/the-governments-code-of-practice-on-copyright-and-ai

[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZVT5N091fb8cciDR9a-t8QAAAFI&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[9] https://www.theguardian.com/australia-news/2023/sep/28/australian-books-training-ai-books3-stolen-pirated

[10] https://www.theregister.com/2023/09/12/openai_copyright_lawsuits/

[11] https://www.theregister.com/2023/11/15/google_deepmind_graphcast/

[12] https://www.theregister.com/2023/11/15/unitedhealthcare_ai_medicine/

[13] https://www.theregister.com/2023/11/14/youtube_ai_label/

[14] https://www.theregister.com/2023/11/14/ai_chemist_paper/

[15] https://www.theregister.com/2023/07/19/meta_llama_2/

[16] https://whitepapers.theregister.com/



Two questions for the price of one

Catkin

It seems that the primary (broader) question is whether ingesting material into a training set, when that material would be legal for a human eye to read or view, is infringement. The secondary question is whether certain materials going into certain training sets were obtained in a way that wouldn't have been legal if a human were reading or viewing them.

For the former, I would propose that pattern extraction is no more copyright infringement than making a graph of how frequently particular words appear in a given book or how many clouds appear in the landscape paintings of a particular artist. In other words, if a human can do it legally, a machine should be able to do it legally; regardless of whether the machine can do it "better" (a subjective appraisal).

Re: Two questions for the price of one

Lurko

Well, certainly in the UK, "copying" includes copying to long term, temporary or random access storage (and thus even caching) of a machine. The fact that a machine may be conceptually "reading" and analysing much as a human would is irrelevant - the machine has made a copy of the original work, and unless anybody wishes to plead altruism on behalf of the LLM owners, does so in most cases for commercial purposes.

Re: Two questions for the price of one

Catkin

I think the law* is currently vague because the language covering copying to and from storage relates to programmes (unless I've missed the part to which you refer, it's a big document). For programs, it's actually quite legal for a legal user to copy, decompile and inspect the operation of programs and no EULA can prohibit this in such a way that renders the actions a copyright infringement (the agreements may break other contracts).

*based on Copyrights, Designs and Patents Act 1988

Re: Two questions for the price of one

mpi

> temporary or random access storage

Like all web browsers do, whenever they access a webpage?

Re: Two questions for the price of one

zuckzuckgo

> whenever they access a webpage?

There is no practical way that any web content can be viewed without that content being temporarily stored in multiple servers on it's route to the viewer. So it would seem to me that if the content is allowed to be published on a publicly accessible website, then they are waving the restriction on temporary storage.

Re: Two questions for the price of one

jdiebdhidbsusbvwbsidnsoskebid

The copyright, designs and patents act in the UK is explicit in that this sort of temporary storage, where only to facilitate use, does not infringe copyright.

Section 28a of the act says:

Copyright in a literary work, other than a computer program or a database, or in a dramatic, musical or artistic work, the typographical arrangement of a published edition, a sound recording or a film, is not infringed by the making of a temporary copy which is transient or incidental, which is an integral and essential part of a technological process and the sole purpose of which is to enable—

(a)a transmission of the work in a network between third parties by an intermediary; or

(b)a lawful use of the work;

and which has no independent economic significance.

Re: Two questions for the price of one

Filippo

> In other words, if a human can do it legally, a machine should be able to do it legally

But that's not obvious at all. There are many things that humans can legally do, but machines can't. Drive a car, in most jurisdictions. Enter legally binding contracts (you can have a machine automate this, but it's the human or company that's bound by it, not the machine). Fart (scale it up, at some point the machine will run into environmental regulations). Memorize large copyrighted texts verbatim (LLMs don't do this, but if they did, they would definitely be infringing; a human wouldn't). More.

I really don't believe that "but it's legal for humans" would be a solid argument in court, no matter how much it may or may not make sense to us.

I'm not saying that LLM training is legal, and I'm not saying that it's illegal. I'm saying that it's not at all clear or obvious which way it goes , existing legislation does not really cover this, and if you attempt to shoehorn this case into current copyright law, I could see very good arguments for it to go either way.

Because of this, I think that the current crop of "AI" is standing on very shaky legal ground, until some high court manages to shoehorn this one way or the other, or lawmakers take action. The idea that someone somewhere wins a case and suddenly the entire industry is illegal is not so far-fetched.

Re: Two questions for the price of one

Catkin

It's just my opinion on the matter, I don't think it applies universally (e.g. driving) but, as far as copyright infringement for deconstructing reality, I think it does. As per your example of memorising, a human that made a perfect reproduction of a copyrighted work from memory by hand would be equally as infringing.

Re: Two questions for the price of one

Doctor Syntax

"The secondary question is whether certain materials going into certain training sets were obtained in a way that wouldn't have been legal if a human were reading or viewing them."

Some books include words to the effect of "not to be stored in an electronic storage system" as a condition of sale. That would be a clear infringement if such a book was used without specific permission. Even if the trained model doesn't contain verbatim text the training data would be an electronically stored copy falling foul of the condition.

As to the wider issue, if the trained model is not a derivative work of all the previous works that were in the training set what's the point of training it on that data as opposed to random lists of words? Would such a trained model be simply fair use of the individual works? My understanding of fair use would be that I can embed one or several quotes from some author(s) into a work which is mostly my own. I'm not sure that embedding the entirety of another's work would count as fair use and much less so the concatenation of several such works in their entirety.

If I were to produce a work which was simply a collection of material from other sources my understanding is that I would have database rights to the collection but not necessarily to the material which went into it. I think I'd have to agree that the training of the model would add database rights for the trainer. However, unless the original material can be passed off as fair use then surely the trained model remains a derivative work of its training material. As such it must surely also include the collected rights of the authors of the training material.

If the production of the derived work is the infringing act then it seems somewhat disingenuous to offer protection against legal costs of those who use a product of it. It's misdirection as to where the potentially infringing act occurred.

Re: Two questions for the price of one

Mike 137

" Some books include words to the effect of "not to be stored in an electronic storage system" as a condition of sale "

The notional get-out for the LLM folks is that the text is not stored, it's scanned and tokenised and the probabilistic relationships between these tokens and all other token so far generated are statistically analysed. It could be argued (and probably will be) that the first stage of this process is essentially no different from borrowing a book and reading it, and once tokenised copyright is not relevant because the text is no longer identifiable from the set of tokens. So it's probably not strictly a derivative work, because the presentation (what is actually protected by copyright) has been entirely eradicated from what is stored.

I don't condone plagiarism (which some might construe this as) but I can't see how it's going to be controlled if the sources used are openly accessible. Offering to cover clients' legal costs for copyright infringement suits only works for those who use the LLM, not the copyright holders (many of whom would not be able to afford the cost of legal redress), and reimbursing copyright holders would involve enormous and complicated human intervention, quite apart from the direct costs.

Probably changing the law to allow the abuse (as is likely here in Blighty in the case of privacy legislation and has just been proposed in the case of the human rights of migrants) will be the ultimate way out.

Knightlie

"We're thieves and parasites profiting from the hard work of others, and we want the law changed to allow us to continue doing that."

Imagine a burglar standing up in court and saying this.

Information wants to be free

Anonymous Coward

I fully understand it might be hard to recover what was exactly taken without permission and from whom.

Perhaps the solution is to make the fruits of this free perpetually. Force the big LLM players to give everyone unlimited access to all their models, data and uses until perpetuity. Force crooks such as Altman and Andreessen to pay any cost in the business from their personal assets to make good until they have to live in a tent and the world is liberated from them.

Re: Information wants to be free

mpi

> Force the big LLM players to give everyone unlimited access to all their models, data and uses until perpetuity

And whos paying for the compute these models require?

Who runs the datacenter where all those GPUs run?

Who pays the people running these datacenters?

Exactly what it's not doing

Mike 137

" It's a large model trained on text data, learning the associations between different ideas Owen Larter

It's not learning about ideas -- it's meaning blind. It's just calculating the statistical associations between tokens the meanings of which it has no comprehension of -- indeed it has no comprehension of anything at all except what's the most likely next token to follow this one. The hype would persuade us that these machines can think, but what they do has no real correspondence with normal human mentation, they're just glorified auto-complete tools.

Legal right to our business model

Missing Semicolon

This is another example of big-business inventing a new way to cheat or steal, using interesting technology, and then to claim that compliance with the law is too hard. So we don't have to, because our business model relies on the mass, uncontrolled theft of something, or other offence.

Examples

Youtube does not moderate meaningfully. Copyrighted stuff stays up, copyright strikes cannot be appealed, creators get taken down for spurious strikes, etc.

Amazon does not check self-published books for plagiarism.

Amazon sells dangerous and knock-off goods, and only takes them down after complaint, it does not pre-emptively check items.

Uber "is not an employer"

Pinterest. Ha!

Real Users find the one combination of bizarre input values that shuts
down the system for days.