News: 1687163409

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Whose line is it anyway, GitHub? Innovation, not litigation, should answer

(2023/06/19)


Opinion Open source. It's open. You can look. Mostly, you can use. There's a clue in the name. Not so fast, claims a class action brought against Microsoft, OpenAI and GitHub. Copilot, an in-IDE AI-powered and open source trained suggestion bot, works by offering lines of code to programmers - and that, the class action suit alleges, breaks the rules, and is being sneaky in trying to hide it. A judge has ruled that some of the claims deserve their day in court. Dear lord, not another copyright battle.

GitHub accused of varying Copilot output to avoid copyright allegations [1]READ MORE

Technology can look very odd to judges. Say you legally purchase an ebook. How do you get it? Routers and caching servers each make copies of the book as it's delivered, but they haven't paid a penny. Are the owners of internet infrastructure breaking copyright billions of times a day? You might think that's a daft question, but it bothered the UK's Supreme Court enough to go to Europe to ask " [2]Is this Internet actually legal ?" Don't be so bloody daft, came the reply. We miss Europe.

How many of the claims against Microsoft, Copilot and OpenAI's code prompter will fall into the bloody daft box remains to be seen. Nobody foresaw AI ingesting global databases of open source code when the rules were written. Then again, nobody foresaw search engines doing wholesale ingestion, analysis and presentation of all the content. That certainly has its problems, but the consensus is that it's too useful and not damaging enough to outlaw. Copilot and other machine learning systems that feed on Internet content are much the same in that respect as search engines. So the question is, is the result not useful enough or too damaging to accept? Where's the balance of interests?

There are useful ways to approach the issues, and they involve - corporate management look away now - ethics. Yes, really, that briefly fashionable chatter about ethical AI offers a concrete way forward that will work a lot better than lawsuits.

Bent out of shape as it is by special interests, the heart of intellectual property law is that the creator's reasonable wishes should be respected. If software is open source, then the creator reasonably wishes people to be able to read it and put them to use. Something that encourages this doesn't seem the worst sin in the world.

[3]

Perhaps it's the way it does it, presenting the code suggestions out of context. There are lots of open source licenses, after all, and some may contain conditions that our happy Copilot cut and paster should know about. Well, assuming Copilot can recognize when it's suggesting someone else's code, it's not unreasonable that it can report the licensing conditions it's offered under. That puts the onus on the coder to comply, which is more ethical than offering up temptation while hiding the consequences. Might even improve the hit rate for following open source rules.

[4]

[5]

What if the original coder really doesn't want their stuff squeezed through the bowels of Copilot? The search engine world tackled that by the invention of robots.txt. Put a file of that name in your web root directory, and you're putting up a "No Entry" sign for web crawlers. Things are a bit more advanced these days, so putting that sort of function into the fabric of GitHub with whatever sort of fine tuning best expresses creator intent would be nice. In any case, telling content providers: "You don't want your stuff in our search results? Fine." has tended to focus minds on ways to live with it. Giving people choices while explaining the consequences? Nice.

Even if giving people the right to remove their code from Copilot and the like results in a ton of good stuff going away, that's not the end of the world. There's the "cleanroom principle", which smashed IBM's dominant position in the 1980s while accelerating the market like crazy. This is something machine learning could learn a lot from.

[6]

The original IBM PC was almost entirely open source. IBM published a technical manual with full circuit diagrams, all using standard chips connected together in standard ways that the chipmakers gave away for free. Designing a functionally equivalent (yet non-copyright) IBM PC clone was something thousands of electronic engineers could do, and hundreds did.

The legal landmine in the beige box was the BIOS, Basic INput-OUtput System, a relatively small chunk of permanent software that provided a standard set of hardware services to operating systems and applications through interrupts - what would be called an API today. If you just copied that code for your clone, IBM would have you bang to rights. You could rewrite the code, but IBM could then tie you up in lawsuits making you prove you didn't copy any of it. Even if you won, the delay and expense would sink you.

[7]Microsoft's Azure mishap betrays an industry blind to a big problem

[8]Windows XP's adventures in the afterlife shows copyright's copywrongs

[9]If you don't get open source's trademark culture, expect bad language

[10]In the battle between Microsoft and Google, LLM is the weapon too deadly to use

Cue the cleanroom. Cloners hired coders who'd never read a line of IBM's BIOS, and forbade them from doing so. These programmers were given the API, which was not copyright, and told to write to that spec. With legal attestations the cloners were happy to swear to in court, the principle that you cannot copy what you haven't seen held - and the last bit of the jigsaw in the original Clone Wars was in place. That APIs provide such a powerful antidote to copyright has led many to try and change their legal status, most recently [11]Google v Oracle . That ended up in the US Supreme Court where it, like all others, failed.

So, take two automated systems, one dedicated to finding and isolating interfaces within code, and one dedicated to applying rules to generate code that provides those interfaces. There's no transfer of lines of code across the virtual air gap. Automated testing of original versus AI code would increase quality. En passant, a very fine set of tools for refactoring would be born, to the benefit of all. Sounds ethical, right?

There we have it. If there are genuine problems with what Copilot is doing, then there are multiple ways to avoid them while preserving utility and creating new benefits. Playing by the rules while making things better? That's a good line to take. ®

Get our [12]Tech Resources



[1] https://www.theregister.com/2023/06/09/github_copilot_lawsuit/

[2] https://www.theregister.com/AMP/2014/06/05/ecj_ruling_prca_v_nla_copyright_ok_on_internet/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZJAnQWR7elhNMdxAJhPYPwAAAM8&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJAnQWR7elhNMdxAJhPYPwAAAM8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJAnQWR7elhNMdxAJhPYPwAAAM8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJAnQWR7elhNMdxAJhPYPwAAAM8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://www.theregister.com/2023/06/12/comment/

[8] https://www.theregister.com/2023/06/05/windows_xp_afterlife/

[9] https://www.theregister.com/2023/04/24/column/

[10] https://www.theregister.com/2023/04/03/opinion_column/

[11] https://www.theregister.com/2021/04/12/oracle_google_case_opinion/

[12] https://whitepapers.theregister.com/



Tom7

Summarising open source as "the creator reasonably wishes people to be able to read it and put them to use" is a vast simplification. Open source licenses have conditions and those conditions are full of minefields for AI ingesting source code. Does use of the code require attribution? How exactly does an AI coding tool comply with that? How does an AI ingesting source code even know that the repo it is ingesting it from belongs to the original author and is correctly following the license terms? If it's been forked within github then it's reasonably straightforward; there are many, many examples of software being copied into github from other places by other people. It's not worth it for the author to go around telling them not to do it, but that doesn't mean the author is okay with it or that they're okay with AI then ingesting it.

NOPENAPI

CrackedNoggin

Interestingly, there are Open AI's user terms that prevent users from building competing AI using Open AI input/output data, for other than research purposes. (See Alpaca AI: Stanford researchers clone ChatGPT AI for just $600). Is, or is that not, exactly the public API interface being described here by our author Paul Kunert? And if so how is it we find that API is now off limits just because NOpen API AI said so?

I'm pretty the any software API's emulated to date (including BIOS) was free to use the original software for comparison in testing - isn't that legally equivalent to using Open AI input/output data in training another LLM model?

The two big open source licenses

Flocke Kroes

Berkley Software Distribution creates a large amount of software and gets paid to do it from government grants. They release code with the BSD license, which is very permissive but includes a requirement to include attribution along with distribution. BSD uses the attribution to show that last year's grant money was put to good use and they should get another grant next year. Several institutions use a BSD-style license - for example MIT. When machine learning software spits out a fragment of BSD style licensed code without attribution and the result gets used then someone has clearly broken the intent of the license. Who that person is and whether it is enough to get them into trouble is something that lawyers will argue about.

The other most common license is GPL. This is about freedom. The authors of GPL software intended to give recipients certain [1]freedoms , but with an important limitation: distributors of GPL software are required to not take those freedoms from the recipients of GPL software. The benefit to the programmer is that the programmer can often get the software they (or their client) wants by making a small change to existing GPL software instead of having to create everything from scratch. When machine learning software spits out GPL code for inclusion into software with a non-GPL compatible license then someone has clearly broken the intent of the license. Who that person is and whether it is enough to get them into trouble is something that lawyers will argue about.

If you think using a snippet of unidentified code supplied by machine learning software will not explode into expensive litigation then I refer you to Oracle vs Google over rangeCheck.

[1] https://en.wikipedia.org/wiki/Free_software#Definition_and_the_Four_Essential_Freedoms_of_Free_Software

The problem is simple: Copyright.

mpi

And I'm not talking about the general principle of copyright.

I am talking about the fact that our current _implementation_ of this principle sucks. Our copyright laws are a relic, fallen out of time. And yes, I am aware that there are almost 200 souvereign nations, each with their own implementations. Doesn't matter, almost all of them suck, and most of them suck big time.

Copyright as it is still implemented, comes from a time when the means of producing AND distributing content were scarce.

That time ended decades ago. Were the laws revised to deal with the fact that everyone could copy things on MCs? Not really. Sure, they were updated, but instead of actually revising if the premises upon which the laws were originally designed still hold true, their "adaptations" simply doubled down on them. Pretend that scarcity in distribution still exists somehow. And if it doesn't, make the law say it does.

What am I talking about? Well, thanks to technology, scarcity in the means of distribution ceased to exist. Yet the laws were designed with that principle in mind. And instead of changing the laws to reflect reality, instead we invented new laws, protecting that inadequate principle, making things ever more absurd. This is how we got to the point where courts have to deal with absurd questions like: is temporarily storing information in RAM to forward it to another client somehow a copyright violation?

And the situation is about to get alot more absurd.

Because we are rapidly approaching a stage where not only distribution, but also _production_ of content will enter a post-scarcity setting. Yes, I am talking about generative AI. We entered an age where people can create an entire novel, even an illustrated one, in a matter of days, with tools that are available for almost everyone. ML systems are cooking up images, text, music, and even short videos. The tech improves so fast one can get dizzy trying to keep up.

How will copyright laws that were designed to deal with a world where Phonograph Cylinders are a hot storage technology, fare in this new reality?

Interesting recap of discussions from commentards[1]

that one in the corner

from earlier articles on Copilot, even if some of the arguments have been pushed a bit far, e.g.

> Even if giving people the right to remove their code from Copilot and the like results in a ton of good stuff going away

Just removing code from Copilot's maw doesn't make it "go away", even Copilot could suggest invoking that code rather than copying it.

Interesting to see if new responses arise here.

[1] which is fine in general, summary articles save a lot of reading, but attribution?!

Reporter: "What would you do if you found a million dollars?"
Yogi Berra: "If the guy was poor, I would give it back."