The truth about Dropbox opening up your files to AI – and the loss of trust in tech
(2023/12/15)
- Reference: 1702601998
- News link: https://www.theregister.co.uk/2023/12/15/dropbox_ai_training/
- Source link:
Comment Cloud storage biz Dropbox spent time on Wednesday trying to clean up a misinformation spill because someone was [1]wrong on the internet.
Through exposure to the social media echo chamber, various people – including Amazon CTO Werner Vogels – became convinced that Dropbox, which introduced a set of [2]AI tools in July, was by default feeding OpenAI, maker of ChatGPT and DALL•E 3, with user files as training fodder for AI models.
Vogels and others [3]advised Dropbox customers to check their settings and opt out of allowing third-party AI services to access their files. For some people, this setting appeared to be opt in; for others, opt out. No explanation was offered by Dropbox.
[4]
Artist Karla Ortiz and celeb Justine Bateman, who like Vogels have significant social media followings, each publicly [5]condemned Dropbox for seemingly automatically, by default, allowing outside AI outfits to drill into people's documents.
[6]
[7]
It was not an implausible scenario, given that tech firms tend to make opt-in the default and OpenAI has refused to disclose its models' training data. The Microsoft-backed machine-learning super lab, for those who haven't been following closely, has been sued by numerous artists, writers, and developers for allegedly training its models on copyrighted content without permission. To date, some of those disputes remain [8]unresolved while others have been [9]thrown out .
While there's widespread outrage among content creators about AI models trained without permission on their work, OpenAI and backers like Microsoft have bet – by offering to [10]indemnify customers using AI services – that they'll prevail in court, or at least make enough money to shrug off potential damages.
[11]
It's a bet that YouTube won. The video sharing site made its name distributing copyrighted clips that its users uploaded. Sued by Viacom for massive copyright infringement in 2007, YouTube escaped liability through the Digital Millennium Copyright Act.
[12]Four more months of Section 702 snooping slipped into $890B US defense budget bill
[13]Google pencils in limited third-party cookie purge for January
[14]Privacy crusaders accuse X of ad-targeting that flouts EU rules
[15]FCC reminds US mobile carriers that customer data needs to be protected
In any event, Dropbox CEO Drew Houston had to set Vogels straight, responding to the Amazonian's post by [16]writing : "Third-party AI services are only used when customers actively engage with Dropbox AI features which themselves are clearly labeled …
"The third-party AI toggle in the settings menu enables or disables access to DBX AI features and functionality. Neither this nor any other setting automatically or passively sends any Dropbox customer data to a third-party AI service."
In other words, the setting is off until a user [17]chooses to integrate an AI service with their account, which then flips the setting on. Switching it off cuts off access to those third-party machine-learning services.
Even so, Houston conceded Dropbox deserved blame for not communicating with its customers more clearly.
[18]
Vogels, however, insisted otherwise. "Drew, this error is completely on me," he [19]wrote . "I was pointed at this by some friends, and with confirmation bias, I drew the wrong conclusion. Instead I should [have] connected with you asking for clarification. My sincere apologies."
Trust gone
That could have been the end of it, but for one thing: as [20]noted by developer Simon Willison, many people no longer trust what big tech or AI entities say. Willison refers to this as the "AI Trust Crisis," and offers a few suggestions that could help – like OpenAI revealing the data it uses for model training. He argues there's a need for greater transparency.
That is a fair diagnosis for what ails the entire industry. The tech titans behind what's been referred to as "Surveillance Capitalism" – Amazon, Google, Meta, data gathering enablers and brokers like Adobe and Oracle, and data-hungry AI firms like OpenAI – have a history of opacity with regard to privacy practices, business practices, and algorithms.
To detail the infractions through years – the privacy scandals, lawsuits, and consent decrees – would take a book. Recall that this is the industry that developed "dark patterns" – ways to manipulate people through interface design – and routinely opts customers into services by default because they know few would bother to make that choice.
Let it suffice to observe that a decade ago Facebook, in a moment of honesty, referred to its [21]Privacy Policy as its [22]Data Use Policy . Privacy has simply never been available to those using popular technology platforms – no matter how often these firms mouth their mantra, "We take privacy very seriously."
Willison concludes that technologists need to earn our trust, and asks how we can help them do that. Transparency is part of the solution – we need to be able to audit the algorithms and data being used. But that has to be accompanied by mutually understood terminology. When a technology provider tells you "We don't sell your data," that isn't supposed to mean "We let third parties you don't know build models or target ads using your data, which remains on our servers and technically isn't sold."
That brings us back to Houston's acknowledgement that "any customer confusion about this is on us, and we'll take a turn to make sure all this is abundantly clear!"
There's a lot of confusion about how code, algorithms, cloud services, and business practices work. And sometimes that's a feature rather than a bug. ®
Get our [23]Tech Resources
[1] https://xkcd.com/386/
[2] https://blog.dropbox.com/topics/product/introducing-AI-powered-tools
[3] https://x.com/Werner/status/1734890651378975007
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[5] https://x.com/kortizart/status/1734862219685528006
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[8] https://www.theregister.com/2023/05/12/github_microsoft_openai_copilot/
[9] https://www.reuters.com/legal/litigation/judge-pares-down-artists-ai-copyright-lawsuit-against-midjourney-stability-ai-2023-10-30/
[10] https://blogs.microsoft.com/on-the-issues/2023/09/07/copilot-copyright-commitment-ai-legal-concerns/
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[12] https://www.theregister.com/2023/12/14/congress_renews_fisa_section_702/
[13] https://www.theregister.com/2023/12/14/google_schedules_limited_thirdparty_cookie/
[14] https://www.theregister.com/2023/12/14/x_illegally_targeted_ads_suit/
[15] https://www.theregister.com/2023/12/13/fcc_sim_swapping_carriers/
[16] https://x.com/drewhouston/status/1735025126671048787?s=20
[17] https://help.dropbox.com/view-edit/privacy-settings-dropbox-ai
[18] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[19] https://twitter.com/Werner/status/1735039349572448364
[20] https://simonwillison.net/2023/Dec/14/ai-trust-crisis/
[21] https://www.facebook.com/privacy/center/
[22] https://about.fb.com/news/2012/05/enhancing-transparency-in-our-data-use-policy/
[23] https://whitepapers.theregister.com/
Through exposure to the social media echo chamber, various people – including Amazon CTO Werner Vogels – became convinced that Dropbox, which introduced a set of [2]AI tools in July, was by default feeding OpenAI, maker of ChatGPT and DALL•E 3, with user files as training fodder for AI models.
Vogels and others [3]advised Dropbox customers to check their settings and opt out of allowing third-party AI services to access their files. For some people, this setting appeared to be opt in; for others, opt out. No explanation was offered by Dropbox.
[4]
Artist Karla Ortiz and celeb Justine Bateman, who like Vogels have significant social media followings, each publicly [5]condemned Dropbox for seemingly automatically, by default, allowing outside AI outfits to drill into people's documents.
[6]
[7]
It was not an implausible scenario, given that tech firms tend to make opt-in the default and OpenAI has refused to disclose its models' training data. The Microsoft-backed machine-learning super lab, for those who haven't been following closely, has been sued by numerous artists, writers, and developers for allegedly training its models on copyrighted content without permission. To date, some of those disputes remain [8]unresolved while others have been [9]thrown out .
While there's widespread outrage among content creators about AI models trained without permission on their work, OpenAI and backers like Microsoft have bet – by offering to [10]indemnify customers using AI services – that they'll prevail in court, or at least make enough money to shrug off potential damages.
[11]
It's a bet that YouTube won. The video sharing site made its name distributing copyrighted clips that its users uploaded. Sued by Viacom for massive copyright infringement in 2007, YouTube escaped liability through the Digital Millennium Copyright Act.
[12]Four more months of Section 702 snooping slipped into $890B US defense budget bill
[13]Google pencils in limited third-party cookie purge for January
[14]Privacy crusaders accuse X of ad-targeting that flouts EU rules
[15]FCC reminds US mobile carriers that customer data needs to be protected
In any event, Dropbox CEO Drew Houston had to set Vogels straight, responding to the Amazonian's post by [16]writing : "Third-party AI services are only used when customers actively engage with Dropbox AI features which themselves are clearly labeled …
"The third-party AI toggle in the settings menu enables or disables access to DBX AI features and functionality. Neither this nor any other setting automatically or passively sends any Dropbox customer data to a third-party AI service."
In other words, the setting is off until a user [17]chooses to integrate an AI service with their account, which then flips the setting on. Switching it off cuts off access to those third-party machine-learning services.
Even so, Houston conceded Dropbox deserved blame for not communicating with its customers more clearly.
[18]
Vogels, however, insisted otherwise. "Drew, this error is completely on me," he [19]wrote . "I was pointed at this by some friends, and with confirmation bias, I drew the wrong conclusion. Instead I should [have] connected with you asking for clarification. My sincere apologies."
Trust gone
That could have been the end of it, but for one thing: as [20]noted by developer Simon Willison, many people no longer trust what big tech or AI entities say. Willison refers to this as the "AI Trust Crisis," and offers a few suggestions that could help – like OpenAI revealing the data it uses for model training. He argues there's a need for greater transparency.
That is a fair diagnosis for what ails the entire industry. The tech titans behind what's been referred to as "Surveillance Capitalism" – Amazon, Google, Meta, data gathering enablers and brokers like Adobe and Oracle, and data-hungry AI firms like OpenAI – have a history of opacity with regard to privacy practices, business practices, and algorithms.
To detail the infractions through years – the privacy scandals, lawsuits, and consent decrees – would take a book. Recall that this is the industry that developed "dark patterns" – ways to manipulate people through interface design – and routinely opts customers into services by default because they know few would bother to make that choice.
Let it suffice to observe that a decade ago Facebook, in a moment of honesty, referred to its [21]Privacy Policy as its [22]Data Use Policy . Privacy has simply never been available to those using popular technology platforms – no matter how often these firms mouth their mantra, "We take privacy very seriously."
Willison concludes that technologists need to earn our trust, and asks how we can help them do that. Transparency is part of the solution – we need to be able to audit the algorithms and data being used. But that has to be accompanied by mutually understood terminology. When a technology provider tells you "We don't sell your data," that isn't supposed to mean "We let third parties you don't know build models or target ads using your data, which remains on our servers and technically isn't sold."
That brings us back to Houston's acknowledgement that "any customer confusion about this is on us, and we'll take a turn to make sure all this is abundantly clear!"
There's a lot of confusion about how code, algorithms, cloud services, and business practices work. And sometimes that's a feature rather than a bug. ®
Get our [23]Tech Resources
[1] https://xkcd.com/386/
[2] https://blog.dropbox.com/topics/product/introducing-AI-powered-tools
[3] https://x.com/Werner/status/1734890651378975007
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[5] https://x.com/kortizart/status/1734862219685528006
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[8] https://www.theregister.com/2023/05/12/github_microsoft_openai_copilot/
[9] https://www.reuters.com/legal/litigation/judge-pares-down-artists-ai-copyright-lawsuit-against-midjourney-stability-ai-2023-10-30/
[10] https://blogs.microsoft.com/on-the-issues/2023/09/07/copilot-copyright-commitment-ai-legal-concerns/
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[12] https://www.theregister.com/2023/12/14/congress_renews_fisa_section_702/
[13] https://www.theregister.com/2023/12/14/google_schedules_limited_thirdparty_cookie/
[14] https://www.theregister.com/2023/12/14/x_illegally_targeted_ads_suit/
[15] https://www.theregister.com/2023/12/13/fcc_sim_swapping_carriers/
[16] https://x.com/drewhouston/status/1735025126671048787?s=20
[17] https://help.dropbox.com/view-edit/privacy-settings-dropbox-ai
[18] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZXvddxEIf6kVi0iAxoPlawAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[19] https://twitter.com/Werner/status/1735039349572448364
[20] https://simonwillison.net/2023/Dec/14/ai-trust-crisis/
[21] https://www.facebook.com/privacy/center/
[22] https://about.fb.com/news/2012/05/enhancing-transparency-in-our-data-use-policy/
[23] https://whitepapers.theregister.com/
Dropbox have been dicks in the past
Gene Cash
People were simply assuming they were continuing to be dicks.
Trust was lost long ago.
IQ test
Cincinnataroo
We know these people lie, we know they don't get much punished, we should expect repeat offences.
If you object to this sort of thing, treat this as an IQ test.
If you don't trust them you can ditch the service (however hard that may be). Test passed.
Anonymous Coward
Look, please stop being so cynical people. Give the guy a break. I personally choose to believe him when he says they aren't selling your data .... this Tuesday (morning)
It's not just a loss of trust.
If you look at what's happening, it's very obviously facts.
Why in the world should the human race put it's trust in a hand full of digital outlaws with no common morals or sense of decency, who have broken all humane laws in the name of self interest and greed time, and time, and time again. These people are very obviously abusing the worst of society against humanity in a self absorbed, and self realized "I'm better than you" abuser game. It's an attempt to feel smarter, more important, and more of a deity by being RICHER, by any means necessary. That very obviously means "hobnobbing", and going to bed with the worst of the worst. These are very bad people, doing very bad things, working against humanity. But, who's going to stop them? Not anyone with any humanity left. The machines won, in the name of quick money. It won't stop there either. It will continue to grow for generations. Destroying all in it's path, long after the monkeys have been assimilated.