Microsoft: Copyright law didn't stop the VCR and shouldn't stop the LLM
- Reference: 1709647294
- News link: https://www.theregister.co.uk/2024/03/05/ms_openai_vs_nyt/
- Source link:
In yesterday's [1]filing [PDF], Microsoft's lawyers recall the early 1980s efforts of the Motion Picture Association to stifle the growth of VCR technology, likening it to the legal efforts of the New York Times (NYT) to stop OpenAI in their work on the "latest profound technological advance."
The motion describes the NYT's allegations that the use of GPT-based products "harms The Times," and "poses a mortal threat to independent journalism" as "doomsday futurology."
[2]
The [3]NYT case is one of many being faced by OpenAI over the training of its Large Language Models (LLMs). The NYT is alleging that large amounts of its content were harvested in that training without permission. It gives examples that it alleges prove ChatGPT was trained using its articles.
[4]
[5]
Microsoft's response doesn't appear to suggest that content has not been lifted. Instead, it says: "Despite The Times's contentions, copyright law is no more an obstacle to the LLM than it was to the VCR (or the player piano, copy machine, personal computer, internet, or search engine.)"
Which seems a bit of a stretch. We're pretty sure Microsoft would be reaching for the phone to its lawyers if bits of Windows were to show up in other operating systems.
[6]
The motion states that the NYT's methods to demonstrate how its content could be regurgitated did not represent real-world usage of the GPT tools at issue. "The Times," explains the motion, "crafted unrealistic prompts to try to coax the GPT-based tools to output snippets of text matching The Times's content."
In its demands for the dismissal of the three claims in particular, the motion points out that Microsoft shouldn't be held liable for end-user copyright infringement through GPT-based tools. It also says that to get the NYT content regurgitated, a user would need to know the "genesis of that content."
"And in any event, the outputs the Complaint cites are not copies of works at all, but mere snippets."
[7]
Finally, the filing delves into the murky world of "fair use," the American copyright law, which is relatively permissive in the US compared to other legal jurisdictions.
OpenAI [8]hit back at the NYT last month and accused the company of paying someone to "hack" ChatGPT in order to persuade it to spit out those irritatingly verbatim copies of NYT content.
[9]OpenAI claims New York Times paid someone to 'hack' ChatGPT
[10]Media experts cry foul over AI's free lunch of copyrighted content
[11]Oracle tells Supremes: Fair use? Pah! There's nothing fair about 'Google's copying'
[12]Reusing software 'interfaces' is fine, Google tells Supreme Court, pleads: Think of the devs
At the time, we said: "By hack, presumably the biz means: Logged in as normal and asked it annoying questions."
The NYT's lead counsel, Ian Crosby, told The Register : "What OpenAI bizarrely mischaracterizes as 'hacking' is simply using OpenAI's products to look for evidence that they stole and reproduced The Times's copyrighted works."
We asked Crosby for his take on Microsoft's motion. We also asked Microsoft how it would react if it found Windows source code being used to train LLMs. We will update this piece should either respond. ®
Get our [13]Tech Resources
[1] https://regmedia.co.uk/2024/03/05/microsoft_motion_to_dismiss_nyt_v_openai_ms.pdf
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZedPt0QwggdJBRC2hUB2XQAAAEY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://www.theregister.com/2023/12/27/the_new_york_times_files/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZedPt0QwggdJBRC2hUB2XQAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZedPt0QwggdJBRC2hUB2XQAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZedPt0QwggdJBRC2hUB2XQAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZedPt0QwggdJBRC2hUB2XQAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[8] https://www.theregister.com/2024/02/28/openai_nyt_lawsuit/
[9] https://www.theregister.com/2024/02/28/openai_nyt_lawsuit/
[10] https://www.theregister.com/2024/01/12/senate_committee_ai_journalism/
[11] https://www.theregister.com/2020/02/13/oracle_v_google/
[12] https://www.theregister.com/2020/01/07/google_vs_oracle_6_jan_supreme_court_brief/
[13] https://whitepapers.theregister.com/
Mere snippets
Does that mean if I install an unlicensed copy of Windows and remove the preponderant mass of unnecessary bloat, I'm not guilty of copyright violation? If so, that, too, is a huge untapped business opportunity.
What an absolute joke of an equivalence
VCRs were personal items that people used to record stuff, and no, copyright law didn't stop the VCR. Though in some countries blank media had a levy to compensate rights owners regardless of what the media was actually going to be used for.
Fast forward a couple of decades to the DVD era. It's now possible to make good quality copies of DVDs, and at scale too. You know what? Copyright law had quite a lot to say about people that did so.
So conflating old analogue tech mostly used by individuals/families with a modern corporate garbage spewer that pilfers other people's content is perhaps the dumbest argument I've yet heard regarding AI.
Re: What an absolute joke of an equivalence
Zactly.
VCRs were largely used to time shift programs for personal use, which is why they passed the court tests.
If a VCR was used to make a new film for commercial use from 100s of snippets of existing films (analogy to LLMs), that would be illegal.
Re: What an absolute joke of an equivalence
Clearly, the more acurate analogy (rather than VCRs) is to view Microsoft as Kim Dotcom, and OpenAI as Megaupload Ltd (or vice-versa).
The VCR was a tool used to do something - and that something could be illegal or not depending on what the user did with it.
LLMs are a tool used to do something - and that something could be illegal or not depending on what the user did with it.
Not only are they the same in that respect, that argument actually means that you still can't use them illegally and still need to get consent for the data you're using, and can't just randomly spew out thousands of copies and sell/give them away without the original owners seeking action against you.
This is a dumb analogy, and actually makes the argument fall against them even worse.
The VCR analogy is spurious, although it does draw attention to the fact that they appear to be duplicating and distributing copyright material without legal authority. Perhaps a better argument in their defence would be that of "derivative works", based on, but substantially different from, the original article. The difference is the part where it gets tricky and ultimately may come to a jury or judge deciding if it is sufficient. Presumably publishers of newspapers, novels and media can implement licencing terms on their website content prohibiting future use in training models? Not a solution for those already ripped off by the LLM models, of course.
"or the player piano"
Please note M$: piano roll manufacturers were eventually obliged to pay royalties.