News: 1687991329

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Microsoft, OpenAI sued for $3B after allegedly trampling privacy with ChatGPT

(2023/06/29)


Microsoft and OpenAI were sued on Wednesday by sixteen pseudonymous individuals who claim the companies' AI products based on ChatGPT collected and divulged their personal information without adequate notice or consent.

The [1]complaint [PDF], filed in federal court in San Francisco, California, alleges the two businesses ignored the legal means of obtaining data for their AI models and chose to gather it without paying for it.

"Despite established protocols for the purchase and use of personal information, Defendants took a different approach: theft," the complaint says. "They systematically scraped 300 billion words from the internet, 'books, articles, websites and posts – including personal information obtained without consent.' OpenAI did so in secret, and without registering as a data broker as it was required to do under applicable law."

[2]

Through their AI products, its claimed, the two companies "collect, store, track, share, and disclose" the personal information of millions of people, including product details, account information, names, contact details, login credentials, emails, payment information, transaction records, browser data, social media information, chat logs, usage data, analytics, cookies, searches, and other online activity.

[3]

[4]

The complaint contends Microsoft and OpenAI have embedded into their AI products the personal information of millions of people, reflecting hobbies, religious beliefs, political views, voting records, social and support group membership, sexual orientations and gender identities, work histories, family photos, friends, and other data arising from online interactions.

OpenAI developed a family of text-generating large language models, which includes GPT-2, GPT-4, and ChatGPT; Microsoft not only champions the technology, but has been [5]cramming it into all corners of its empire, from Windows to Azure.

[6]

"With respect to personally identifiable information, defendants fail sufficiently to filter it out of the training models, putting millions at risk of having that information disclosed on prompt or otherwise to strangers around the world," the complaint says, citing The Register 's [7]March 18, 2021 special report on the subject.

The 157 page complaint is heavy on media and academic citations expressing alarm about AI models and ethics but light on specific instances of harm.

For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model.

[8]Microsoft Azure OpenAI lets enterprises feed corporate secrets to ChatGPT

[9]OpenAI calls for tough regulation of AI while quietly seeking less of it

[10]Open source licenses need to leave the 1980s and evolve to deal with AI

[11]Small custom AI models are cheap to train and can keep data private, says startup

It remains to be seen how, if at all, plaintiff-created content and metadata has actually been exploited and whether ChatGPT or other models will reproduce that data.

OpenAI in the past has dealt with the reproduction of personal information [12]by filtering it .

[13]

The lawsuit is seeking class-action certification and damages of $3 billion – though that figure is presumably a placeholder. Any actual damages would be determined if the plaintiffs prevail, based on the findings of the court.

The complaint alleges Microsoft and OpenAI have violated America's Electronic Privacy Communications Act by obtaining and using private information, and by unlawfully intercepting communications between users and third-party services via integrations with ChatGPT and similar products.

The sueball further contends the defendants have violated the Computer Fraud and Abuse Act by intercepting interaction data via plugins.

It also alleges violations of the California Invasion of Privacy Act and unfair competition law, the Illinois Biometric Information Privacy Act and consumer fraud and deceptive business practices law, and New York business law, along with various general harms (torts) like negligence and unjust enrichment.

Microsoft and OpenAI declined to comment.

Microsoft, its GitHub subsidiary, and OpenAI were sued last November for allegedly reproducing the code of millions of software developers in violation of licensing requirements through the Copilot service, based on an OpenAI model, that GitHub offers. That case [14]is ongoing . ®

Get our [15]Tech Resources



[1] https://storage.courtlistener.com/recap/gov.uscourts.cand.414754/gov.uscourts.cand.414754.1.0.pdf

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZJ0B44D4kp@FpBt45W2g9gAAANc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJ0B44D4kp@FpBt45W2g9gAAANc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJ0B44D4kp@FpBt45W2g9gAAANc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://www.theregister.com/2023/05/10/microsoft_copilot_ai/

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJ0B44D4kp@FpBt45W2g9gAAANc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://www.theregister.com/2021/03/18/openai_gpt3_data/

[8] https://www.theregister.com/2023/06/22/microsoft_azure_ai_data/

[9] https://www.theregister.com/2023/06/21/openai_government_regulation/

[10] https://www.theregister.com/2023/06/23/open_source_licenses_ai/

[11] https://www.theregister.com/2023/06/22/small_custom_ai_models/

[12] https://www.theregister.com/2021/03/18/openai_gpt3_data

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJ0B44D4kp@FpBt45W2g9gAAANc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[14] https://www.theregister.com/2023/06/09/github_copilot_lawsuit/

[15] https://whitepapers.theregister.com/



Should sue them for...

druck

...the full $10bn invested in OpenAI, and every penny of profit from AI based on data scraped without permission.

TheMaskedMan

"They systematically scraped 300 billion words from the internet, 'books, articles, websites and posts – including personal information obtained without consent."

Who is alleged to have obtained the information without consent? OpenAI, or the sites they scraped?

Either way, if you're going to make information about yourself available on websites and posts, you're not too concerned about who sees it, are you? Or, if you are, and you still make it available, you're too bloody stupid to be allowed access to the net.

Material in books and articles is a little more tricky, in that the subject may not have been - and probably wasn't - the author, and may not have had much say in what was written. Even so, the material is published and available to anyone who wants to find it; if the subject doesn't like what was written, they need to take that up with the author.

The situation would be vastly different if the complainants had give their data to OpenAI directly, and possibly in confidence. But to complain about them using information that is already publicly available strikes me as pure opportunism. Are the same people suing Google for indexing the pages containing their data? Or the sites that have been scraped? If not, why not?

Of course, Google has long had a policy of (eventually) removing, or at least hiding, some personal information on request. Have the complainants approached OpenAI and asked them to remove their data? Again, if not, why not?

The likelihood of anyone having suffered any real harm from the likes of chatGPT seems to be pretty remote to me. I recall hearing some statistics recently (can't recall where, unfortunately) that suggested that a very large percentage of Americans had never heard of chatGPT, and of those that had heard of it, most hadn't used it. What are the chances that the relatively small proportion of people who have used it did so with a view to asking it about the complainants' personal data, rather than getting it to spew out semi-reliable filler for their homework / website / court documents etc? Pretty much zero, I should think. But who cares about that that when there's potentially money to be made from litigation?

doublelayer

And yet, if I went to a dump of data which contained data about you and made it a lot more public, you'd still have objections and I would still be breaking the law. It does not matter that I didn't steal it in the first place, nor does it matter how the original source got the data (if they stole it or if you gave it voluntarily, you did not authorize its publication). This only applies to certain types of personal information, and the particular jurisdiction will determine whether some information gets protection or not. The people in this case are complaining about increased publication of their details, and their case will succeed or fail based on that argument.

Even if this was only about collecting posts they have voluntarily written and published, they could still attempt to block OpenAI from repeating it using copyright claims. Just because something can be accessed does not mean you have the right to distribute it. Laws exist which limit your rights in that area, both in privacy and elsewhere. In practice, people should be careful not to publish anything they wouldn't like to see abused, because a lot of people will not obey those laws, but just because that will happen doesn't change the fact that they have rights over some of that data and they have a chance of getting a court to penalize those who violate them.

Private Information About Yourself Made Publicly Available

An_Old_Dog

Either way, if you're going to make information about yourself available on websites and posts, you're not too concerned about who sees it, are you?

There's a substantial difference between knowingly posting information about yourself versus having it sucked out of your browser by some website's Javascript, because that one freakin' time you forgot to clear all your cookies before clicking a link from one website to a second website.

To successfully defend your personal information you have to never screw up; the attacking websites have to succeed only once, and then they've got your info.

<FrikaC> I should probably reboot...
<FrikaC> ok brb
<FrikaC> So, what apart form avoiding virii, memory leaks, and rampant
crashing does Linux reallhy offer :)
<LordHavoc> reliable multitasking?