Microsoft calls AI privacy complaint 'doomsday hyperbole'
- Reference: 1710235866
- News link: https://www.theregister.co.uk/2024/03/12/microsoft_doomsday_hyperbole_ai_filing/
- Source link:
A variation on that phrase, which appeared in the Windows giant's [1]motion to dismiss [PDF] the privacy lawsuit, surfaced in early March in a similar [2]court filing [PDF] supporting a motion to dismiss an AI copyright claim brought by The New York Times.
In the copyright case, Redmond's legal team referred to the news organization's enumeration of AI harms as "doomsday futurology."
[3]
Through their attorneys, the thirteen plaintiffs in the September 2023 [4]privacy claim [PDF], T. et al v. OpenAI LP et al (3:23-cv-04557-VC), last week submitted responses opposing the motions to dismiss filed by both OpenAI and Microsoft.
[5]
[6]
An event akin to doomsday could await OpenAI and Microsoft if the plaintiffs prevail: they seek relief including an injunction that would force the defendants to let people remove their data from AI models.
[7]Filing NeMo: Nvidia's AI framework hit with copyright lawsuit
[8]AI models show racial bias based on written dialect, researchers find
[9]You got legal trouble? Better call SauLM-7B
[10]Microsoft: Copyright law didn't stop the VCR and shouldn't stop the LLM
The crux of the complaint is that OpenAI and Microsoft allegedly trained their models on data scraped from the web without first securing adequate consent, and now continue to harvest personal information through API integrations with their products.
The privacy complaint, overseen by Morgan and Morgan Complex Litigation Group and Clarkson Law Firm, accuses the defendant companies of failing to filter personal information out of their training models and "putting millions at risk of having that information disclosed on prompt or otherwise to strangers around the world." It cites, among other sources, a [11]Register article to support its claims.
The legal filing goes on to insist that the developers' API-based data harvesting includes "user locations and image-related data obtained through Snapchat, user financial information through Stripe, musical tastes and preferences through Spotify, user patterns and private conversation analysis through Slack and Microsoft Teams, and even private health information obtained through the management of patient portals such as MyChart."
[12]
Microsoft, in its motion to dismiss, argues "Plaintiffs do not plead any facts plausibly showing they have been affected by any of the supposed 'scraping,' 'intercepting,' and 'eavesdropping' they allege. Nowhere do they say what of their private information Microsoft ever improperly collected or used; nor do they identify any harm they individually suffered from anything that Microsoft allegedly did."
The software giant contends that the plaintiffs have not stated a valid claim.
OpenAI likewise argues the plaintiffs have not sufficiently elaborated what personal information was allegedly stolen. The AI biz also maintains that people who use its products have agreed to the terms of use. "Further, Plaintiffs' novel theory that companies cannot use publicly available online information, or information provided by their own users, to train and improve their products is legally baseless, and none of the 11 asserted claims provide a remedy for such conduct," OpenAI's [13]motion to dismiss [PDF] reads.
[14]
The plaintiffs, in their efforts to convince the judge to allow their claim to proceed, argue that the way these companies have handled AI is simply wrong.
The motion opposing OpenAI's argument declares:
"OpenAI gave no notice to the world that, for years, it was secretly harvesting from the internet everything ever created and shared online, anywhere, by hundreds of millions of Americans."
"That, for a decade plus, every consumer's use of the internet thus operated as a gratuitous donation to OpenAI: of our insights, talents, artwork, personally identifiable information, copyrighted works, photographs of our families and children, and all other expressions of our personhood – for products that stand to concentrate the country's wealth in even fewer corporate behemoths, displace jobs at scale, and risk the future of mission-critical industries like art, music, and journalism, while creating dangerous new industries like the high-speed spawning of child pornography. It is no wonder the public is outraged by the largest-ever theft of data – to which no one consented."
And the plaintiffs' argument challenging Microsoft's motion covers similar ground.
"Microsoft’s motion essentially points to its strategic business partner OpenAI's motion to dismiss and says 'we agree.' That is hardly surprising. But its motion fails for the same reasons as its partner-in-theft. Plaintiffs' allegations exceed applicable pleading standards, stating legal claims supported factually by hundreds of sources. The only way the complaint could be more specific factually would be if Microsoft and OpenAI would once and for all open the 'black box' of training data – that they won't let anyone see."
The lawsuit alleges violations of the Electronic Communications Privacy Act, the Comprehensive Computer Data Access And Fraud Act, the California Invasion of Privacy Act, and various California and Illinois competition and privacy laws.
Ryan Clarkson, managing partner of Clarkson Law Firm, told The Register in an email, "Practically, OpenAI's legal position would forever change the internet, in that the only way not to surrender all of our personal information, family photographs, copyrighted works, art, and more would be to cease using the internet altogether."
"Fortunately the law compels a different result: choice and compensation for the millions of Americans who never consented to the mass theft of their personal information without which OpenAI's $100 billion dollar business would be worth zero."
Judge Vince Chhabria's decision on the motions to dismiss – whenever it happens – will determine whether the privacy claim can continue. ®
Get our [15]Tech Resources
[1] https://storage.courtlistener.com/recap/gov.uscourts.cand.417880/gov.uscourts.cand.417880.53.0.pdf
[2] https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.65.0.pdf
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offbeat/legal&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZfA1yRmsAApIVysfyA8@ogAAAYA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://storage.courtlistener.com/recap/gov.uscourts.cand.417880/gov.uscourts.cand.417880.45.0.pdf
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offbeat/legal&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZfA1yRmsAApIVysfyA8@ogAAAYA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offbeat/legal&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZfA1yRmsAApIVysfyA8@ogAAAYA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2024/03/11/authors_file_lawsuit_to_torpedo/
[8] https://www.theregister.com/2024/03/11/ai_models_exhibit_racism_based/
[9] https://www.theregister.com/2024/03/09/better_call_saul_llm/
[10] https://www.theregister.com/2024/03/05/ms_openai_vs_nyt/
[11] https://www.theregister.com/2021/03/18/openai_gpt3_data/
[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offbeat/legal&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZfA1yRmsAApIVysfyA8@ogAAAYA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[13] https://storage.courtlistener.com/recap/gov.uscourts.cand.417880/gov.uscourts.cand.417880.50.0_1.pdf
[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offbeat/legal&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZfA1yRmsAApIVysfyA8@ogAAAYA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[15] https://whitepapers.theregister.com/
Not surprising
Borkzilla automatically brands any privacy concern as doomsday fodder, because truly respecting our privacy would indeed be doomsday for it and many others.
the only way...
...not to surrender all of our personal information, family photographs, copyrighted works, art, and more would be to cease using the internet altogether.
Um... yes? Good. You finally get it. Were they not paying attention for the last 25 years? The Internet is a public space, and where it's not public, it's a space where you are permitted to exist by the companies that own it. It always has been this way. Unless you run your own server and set your own terms about who can access it there has never been any expectation of privacy here. We were telling people 25 years ago that you need to be careful what you say online and who you say it to because once something is "out there" it's basically impossible to get it back. This isn't new, it's just something that - for lack of a better term "regular" - people are waking up to now because it's in the news. We've known this since the beginning. The internet isn't safe. It's never been safe.
Considering at one point you could prompt Chat GPT to generate Windows activation keys which were obviously scraped from publicly available sources, but then as soon as MS got wind of that they shut down that function to replace it with a notice on how piracy = bad.
So MS clearly don't like their own works being in Open AI's LLM if it might loose Microsoft money, but are happy with other peoples work being scraped without permission.
Everything but everything is hyperbole
with the AI hype train of venture capital bullshit.
Why the US doesn't get privacy episode 239,203,829,178
MS argued:
Plaintiffs do not plead any facts plausibly showing they have been affected by any of the supposed 'scraping,' 'intercepting,' and 'eavesdropping' they allege.
As this is the US and there is no concept of privacy, you, the little guy, have to somehow prove monetary loss because your individual items of data were scraped, but big tech is allowed to scrape everything and use it all to make a new product which brings in billions.
Where "little guy" means any person or business with a market cap smaller than Microsoft's.
I don't say they're wrong to bring the action but it would be a good idea to do a bit of prompt engineering to prove that the data can be regurgitated. Courts like evidence.
Nowhere do they say what of their private information Microsoft ever improperly collected
I thought that they had proof of the AI spitting out NYT stories verbatim? If that's so, then MS are claiming that if it's on a public website, it's not "private" and can therefore be scraped and regurgitated, despite the copyright notices on that public website. If that's the case, the NYT could create a Large Operating System model, and scrape copies of Windows from the public Microsoft site and then regurgitate them too.