Open source licenses need to leave the 1980s and evolve to deal with AI
- Reference: 1687509013
- News link: https://www.theregister.co.uk/2023/06/23/open_source_licenses_ai/
- Source link:
[1]AI was born from open source software . But the free software and open source licenses, based on copyright law, to deal with software code are not a good fit for the large language model (LLM) neural nets and datasets that fuel AI's open source software. Since many programming datasets, in particular, are based on free software and open source code, something must be done. And that's why Stefano Maffulli, [2]Open Source Initiative (OSI) executive director, and a host of other open source and AI leaders are working on combining AI and open source licenses in ways that will make sense for both.
Lest you think this is some kind of theoretical, legal discussion with no impact on the real world, think again. Consider [3]J. Doe 1 et al vs GitHub . The plaintiffs in this case in the United States Northern District Court of California allege Microsoft, OpenAI, and GitHub, via their commercial AI-based system, OpenAI's Codex and GitHub's Copilot, had [4]ripped off their open source code . The result? The plaintiffs claim that "suggested" code consists of often near-identical copies of code scraped from public GitHub repositories, without the required open source license attributions.
[5]
This [6]case continues . The [7]amended complaint includes accusations of violating the Digital Millennium Copyright Act, breach of contract (open source license violations), unfair enrichment, and unfair competition claims, and breach of contract (selling licensed materials in violation of GitHub's policies).
[8]
[9]
Don't think this kind of lawsuit is just Microsoft's problem. It's not. Sean O'Brien, a Yale Law School lecturer in cybersecurity and founder of the [10]Yale Privacy Lab , told my colleague David Gewirtz: "I believe there will soon be an [11]entire sub-industry of trolling that mirrors patent trolls , but this time surrounding AI-generated works. A feedback loop is created as more authors use AI-powered tools to ship code under proprietary licenses. Software ecosystems will be polluted with proprietary code that will be the subject of cease-and-desist claims by enterprising firms."
He's right. I've been covering patent trolls for decades. I guarantee that licensing trolls will come after "your" ChatGPT and Copilot code.
[12]
Some people, such as Felix Reda, a German researcher and politician, claim that all [13]AI-produced code is public domain . US attorney [14]Richard Santalesa , a founding member of the [15]SmartEdgeLaw Group , observed to Gewirtz that there are contract and copyright law issues. They're not the same thing. Santalesa believes companies producing AI-generated code will "as with all of their other IP, deem their provided materials – including AI-generated code – as their property." In any case, however, [16]public domain code is not the same thing as open source code .
[17]Will Flatpak and Snap replace desktop Linux native apps?
[18]Red Hat promises AI trained on 'curated' and 'domain-specific' data
[19]EU's Cyber Resilience Act contains a poison pill for open source developers
[20]Here's how the data we feed AI determines the results
On top of all that, there's the whole issue of how the datasets should be licensed. There are [21]many "open" datasets under numerous open source licenses, but it's not usually a good fit.
In our conversation, Open Source Initiative's Maffulli elaborated on how various artifacts produced by AI and machine learning systems fall under different laws and regulations. The open source community must determine which laws best serve their interests. Maffulli compared the current situation to the late '70s and '80s when software emerged as a distinct discipline, and copyright began to be applied to the source and binary codes.
We're at a similar crossroads today. AI programs such as TensorFlow, PyTorch, and Hugging Face Hub work well under their open source licenses. The new AI artifacts are another story. Datasets, models, weights, etc. don't fit squarely into the traditional copyright model. Maffulli argued that the tech community should devise something new that aligns better with our objectives, rather than relying on "hacks."
Specifically, open source licenses designed for software, Maffulli noted, might not be the best fit for AI artifacts. For instance, while MIT License's broad freedoms could potentially apply to a model, questions arise for more complex licenses like Apache or the GPL. Maffulli also addressed the challenges of applying open source principles to sensitive fields like healthcare, where regulations around data access pose unique hurdles. The short version of this is that medical data can't be open sourced.
[22]
Simultaneously, most commercial LLMs datasets are black boxes. We literally don't know what's in them. So we end up, as the Electronic Frontier Foundation (EFF) puts it, in a situation where we have [23]"Garbage In, Gospel Out." We need, the EFF concludes, open data.
So it is that the OSI, said Maffulli, together with Open Forum Europe, Creative Commons, Wikimedia Foundation, Hugging Face, GitHub, the Linux Foundation, ACLU Mozilla, and the Internet Archive are working on a draft for defining a common understanding of open source AI principles. This will be "critical in conversations with legislative bodies." Even now, EU, US, and UK government agencies are struggling to develop AI regulation, and they're woefully under-equipped to deal with the issues.
Stefano concluded by saying we should start with "a return to the basics," the [24]GNU Manifesto , which predates most licenses and sets the "North Star" for the open source movement. Maffulli suggested that its principles remain surprisingly relevant when applied to AI systems. By focusing on first principles, we'll be better able to navigate this complex intersection of AI and open source. ®
Get our [25]Tech Resources
[1] https://www.theregister.com/2023/03/24/column/
[2] https://opensource.org/
[3] https://dockets.justia.com/docket/california/candce/4:2022cv06823/403220
[4] https://www.theregister.com/2022/10/19/github_copilot_copyright/
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZJVtRWR7elhNMdxAJhNQSAAAAM4&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[6] https://www.theregister.com/2023/05/12/github_microsoft_openai_copilot/
[7] https://www.theregister.com/2023/06/09/github_copilot_lawsuit/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJVtRWR7elhNMdxAJhNQSAAAAM4&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJVtRWR7elhNMdxAJhNQSAAAAM4&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] https://privacylab.yale.edu/
[11] https://www.zdnet.com/article/if-you-use-ai-generated-code-whats-your-liability-exposure/
[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJVtRWR7elhNMdxAJhNQSAAAAM4&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[13] https://felixreda.eu/2021/07/github-copilot-is-not-infringing-your-copyright/
[14] https://twitter.com/RichNet
[15] http://www.smartedgelawgroup.com/
[16] https://opensource.com/article/19/10/shareware-vs-open-source
[17] https://www.theregister.com/2023/06/09/will_flatpak_and_snap_replace/
[18] https://www.theregister.com/2023/05/26/red_hat_ai/
[19] https://www.theregister.com/2023/05/12/eu_cyber_resilience_act/
[20] https://www.theregister.com/2023/04/28/column/
[21] https://github.com/eugeneyan/open-llms
[22] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJVtRWR7elhNMdxAJhNQSAAAAM4&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[23] https://www.eff.org/deeplinks/2023/01/open-data-and-ai-black-box
[24] https://www.gnu.org/gnu/manifesto.en.html
[25] https://whitepapers.theregister.com/
If you were fully aware of what's in a lot of the training data, you'd probably agree with the standard definition.
Unsettle law
At the moment, it is not clear what the law is or even should be. Providing an ML code generating service is a legal risk. Using an ML code generating service is a risk. I would vote for:
*) GPL in, GPL out
*) BSD in, BSD out
From then on it gets more difficult:
*) BSD+MIT in, who gets attribution?
*) Many licences in, output code cannot safely be distributed under any license.
Re: Unsettle law
"who gets attribution?"
I can point to software that states upfront "This software or parts of this software are provided under one or more of the following licenses ...".
"Providing an ML code generating service is a legal risk. Using an ML code generating service is a risk."
Yes. And the lawyers are eventually going to figure that out. Gut feeling is we're heading for another AI Winter based on this aspect alone.
Let us hope that we can ignore licensing
This seems to be the attitude. Grab/digest as much code as possible, learn from it, generate proprietary code as a result.
Having to worry about the license of the code digested is hard but should be doable. It will increase costs and this is what the ML owners do not want.
Copying of code and ignoring licenses has been going on since the year dot. If the output code is not distributed then it is hard to detect. So what is happening here is not new but just happens faster.
I don't think open-source licenses need to change at all. There have *always* been criminals that have tried to exploit it without adhering to the license. This is not a new thing.
Companies (including AI companies) are just waking up to this wealth of functionality that has been slowly and steadily growing (whilst the proprietary stuff has burned out due to corporate lifespan or over monetization). The responsibility is on them to not break license terms. But this is what companies do; they push legal boundaries for maximum wealth.
If anything, what the open-source movement needs is a way of *mass* detecting when i.e GNU licensed code has been used and perhaps come up with a concept of crowd funded lawyers to tangle up the company in breach. Perhaps we should also consider AI GPL lawyers to churn through all the many companies I am sure are in breach.
How far do you take it?
How of you cope with code fragments that could be "original" or found in many projects?
for ( int i = 0; i
Even if it represents an algorithm to do "X", there could be many implementations of the same algorithm out there.
Some crazy software patent attempts have been made along the lines of "using a loop to search for...".
I'd argue the current system is "Gospel In, Garbage Out" .