4chan and other web sewers scraped up into Google's mega-library for training ML
- Reference: 1681975813
- News link: https://www.theregister.co.uk/2023/04/20/google_c4_data_nasty_sources/
- Source link:
An [1]investigation by The Washington Post and the Allen Institute for AI analyzed Google's immense public [2]C4 dataset , released for academic research, to get a better understanding of what types of websites are typically scraped to train large language models.
The C4 dataset was used to train Google's T5 Text-to-Text Transfer Transformer as well as Facebook's Large Language Model Meta AI (LLaMA), a variant of which [3]raised alarm bells .
[4]
It appears C4 has ingested concerning material, which is being used to build next-gen machine-learning systems. That potentially could cause those systems to behave inappropriately and unreliably.
[5]
[6]
Regular Register readers will be aware we've pointed out problems with training datasets over and over, such as the horrible underbelly of a highly cited set [7]curated by MIT .
Latest probe
The Post and Allen Institute's analysts ranked the top 10 million websites included in C4 by matching text that appeared as internet content. Although C4 is a smaller, cleaner version of the Common Crawl dataset, which comprises text from billions of websites, it still contained undesirable material from dark corners of the internet.
Racist, anti-trans, and toxic text were scraped from websites such as the race-hate haven Stormfront, the doxxing forum Kiwi Farms, and toxic message board 4chan. It's therefore unsurprising that language models based on that corpus can generate inappropriate content, talk of conspiracy theories, or bring up dubious ideologies.
C4 is also made up of websites hosting degrees of personal information, such as voter registration databases. In the background of this, several regulatory agencies in Italy, Canada, Spain, and France have since launched investigations into OpenAI's ChatGPT over data privacy concerns, since the model can ingest and generate sensitive information.
[8]
Large language models powering AI chatbots are not intelligent nor conscious, no matter how magic they seem: they write by predicting the flow of words and sentences in response to prompts, questions, and instructions from users or even other bots. This involves drawing upon the mountains of data they've been trained on, and learning from it, to emulate what a person would write.
These predictions therefore reflect patterns in the kinds of text humanity produces, such as internet posts, news articles, poetry, and novels, that is all vacuumed up into vast training datasets.
These systems cannot tell fact from fiction, are fed vast amounts of data scraped from the internet, and can generate inaccurate results as well as regurgitate information.
[9]
Companies that build large language models try to filter out unwanted content, in the training and inference stages, though their review processes are imperfect. What's also frustrating is that builders of commercial AI models - such as OpenAI's ChatGPT, Microsoft's new Bing, or Google's Bard chat - don't always disclose how they sourced, scrubbed, and processed their training data.
[10]Reddit: If you want to slurp our API to train that LLM, you better pay for it, pal
[11]Predict stocks, foresee public opinion, all kinda possible with ChatGPT-like models
[12]Google crams more AI into search as Apple, Samsung sniff around Bing
[13]What if someone mixed The Sims with ChatGPT bots? It would look like this
Fortunately, the C4 dataset isn't as bad as others: it mostly contains material scraped from more benign websites spanning journalism, software development, medicine, and content creation. Most of its text comes from Google patents, Wikipedia, and Scribd. The New York Times and scientific journals from academic publisher PLOS ranked fourth and fifth respectively by volume in the dataset. C4 also features content from individuals' blogs, religious websites, and more.
Copyrighted material is swept up in the dataset, too, with the © symbol appearing more than 200 million times. It's not clear whether companies building AI products based on training data containing protected works are liable for infringing intellectual property.
Stability AI, a startup building text-to-image tools has been sued for scraping copyrighted images from stock photo platforms. OpenAI also faces a lawsuit challenging its collection of public code hosted on GitHub used to create Microsoft's AI-pair-programming Copilot tool.
Reddit just [14]announced an update to its terms and conditions for its API services, requiring companies to pay for licenses to scrape its data. "We are introducing a new premium access point for third parties who require additional capabilities, higher usage limits, and broader usage rights," it stated on Tuesday.
C4 contains content from the internet up until 2019, but as other more recent models were built with similar data collection practices this research shines a light on how AI chatbots can produce problematic output.
The Register has asked the Allen Institute of AI for further comment. ®
Get our [15]Tech Resources
[1] https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/
[2] https://www.tensorflow.org/datasets/catalog/c4
[3] https://www.theregister.com/2023/03/21/stanford_ai_alpaca_taken_offline/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZEENQnG3zr-sv-EVcdCSwgAAABU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZEENQnG3zr-sv-EVcdCSwgAAABU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZEENQnG3zr-sv-EVcdCSwgAAABU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2020/07/01/mit_dataset_removed/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZEENQnG3zr-sv-EVcdCSwgAAABU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZEENQnG3zr-sv-EVcdCSwgAAABU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] https://www.theregister.com/2023/04/18/reddit_charging_ai_api/
[11] https://www.theregister.com/2023/04/18/large_language_models_like_chatgpt/
[12] https://www.theregister.com/2023/04/17/google_bing_ai_race/
[13] https://www.theregister.com/2023/04/11/sims_ai_generation/
[14] https://www.theregister.com/2023/04/18/reddit_charging_ai_api/
[15] https://whitepapers.theregister.com/
Devil’s advocate
If these toxic sewers are written and populated by meat-bags, why exclude them from training data?
Surely everyone is entitled to be heard, no?
Re: Devil’s advocate
>Surely everyone is entitled to be heard, no?
If we want AI to behave like humans- absolutely. That's humans' advocate!
But is that what we want?
I'd prefer an AI trained on the scientific principle: suggest explanations from the evidence, test the sh't out of them, drop the fails, and keep on testing. The last one standing- go with that for now.
There's an AI for the future.
OK OK, godamnit- Musk got there first... I *really hate* that guy...
Re: Devil’s advocate
As the article points out, it doesn't understand the data... it's just 01010101010... it does not (yet) seem to understand what's good or bad, just that something matches with your input
It's still a small child and it definitely needs a spanking until it stops peeking at Daddy's
I don't mind 4chan being used. It is the use of El Reg comments that scares the shit out of me.
(c) Rune
Ah, yes. 4chan.
The kiddies who trolled QAnon into existence "for the lulz". Just the folks to feed into the AI of your choice.
Think Redmond had trouble with Tay? Might want to stock up on popcorn. I'll bring the beer.
The idea of a chatbot being trained on comments from amanfrommars1 amuses me. As above so below, and all that.
Wasn't there an expression mooted oh, many, many years sinceupon?
"Garbage in, garbage out."
Seems very, very relevant to AI text generators.
Exactly this. Data science products, including AIs, are only as good as the data sets that they’re trained on. Clean data set? Good output. Dirty dataset? Bad output.
If our AIs are going to be truly useful then they need to be trained on scientific principles. They need to learn based on evidence and accept that they might be wrong. Therefore, they need to be trained from the best materials available. Reputable and peer reviewed scientific journals, historical documents, literature - all of it taken from around the world, not just one specific region, and they need to be retrained regularly.
What they shouldn’t be trained with is data scraped from the whole of the internet. Do that and you end up with an argumentative, prejudiced, partisan AI. Not actually something which is of particular use to humanity as a whole.
"Problematic, racist, and pornographic web content"
I understand that racist web content should not be used in training data. I'm a bit less sure about pornographic web content, but I'll give that one a pass.
Now, if web content is neither racist nor pornographic, how exactly is it "problematic". What is the definition of "problematic" in that context ?
Could someone enlighten me ?
Re: "Problematic, racist, and pornographic web content"
"how exactly is it "problematic""
It doesn't agree with their shaman of choice would be my guess.
Bearing in mind some of the data sources, how long before the AI decides that humanity is something evil and needs to be stamped out?
Elon Musk has said on Twitter he'll sue anyone that's been using Twitter's data for AI training. This sounds like it might be fun - popcorn is already in the microwave.
To answer the question ...
"Are you still so keen to have generative AI write your emails, sales proposals, blog posts ... ?"
No, I am not. But then, being both educated and sane, I never was.
Thank you for asking.