News: 1654754229

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

AI chatbot trained on posts from web sewer 4chan behaved badly – just like human members

(2022/06/09)


A prankster researcher has trained an AI chatbot on over 134 million posts to notoriously freewheeling internet forum 4chan, then set it live on the site before it was swiftly banned.

Yannic Kilcher, an [1]AI researcher who posts some of his work to YouTube, called his creation "GPT-4chan" and [2]described it as "the worst AI ever". He trained GPT-J 6B, an open source language model, on a [3]dataset containing 3.5 years' worth of posts scraped from 4chan's imageboard. Kilcher then developed a chatbot that processed 4chan posts as inputs and generated text outputs, automatically commenting in numerous threads.

Netizens quickly noticed a 4chan account was posting suspiciously frequently, and began speculating whether it was a bot.

[4]

4chan is a weird, dark corner of the internet, where anyone can talk and share anything they want as long as it's not illegal. Conversations on the site's many message boards are often very odd indeed – it can be tricky to tell whether there is any intelligence, natural or artificial, behind the keyboard.

[5]

[6]

GPT-4chan behaved just like 4chan users, spewing insults and conspiracy theories before it was banned.

The Reg tested the model on some sample prompts, and got responses ranging from the silly and political to offensive and anti-Semitic.

[7]

It probably didn't do any harm posting in what is already a very hostile environment, but many criticized Kilcher for uploading his model. "I disagree with the [8]statement that what I did on 4chan, letting my bot post for a brief time, was deeply awful (both bots and very bad language are completely expected on that website) or that it was deeply irresponsible to not consult an institutional ethics review board," he told The Register .

"I don't disagree that research on human subjects is not to be taken lightly, but this was a small prank on a forum that is filled with already toxic speech and controversial opinions, and everybody there fully expects this, and framing this as me completely disregarding all ethical standards is just something that can be flung at me and something where people can grandstand."

Kilcher did not release the code to turn the model into a bot, and said it would be difficult to repurpose his code to create a spam account on another platform like Twitter, where it would be riskier and potentially more harmful. There are several safeguards in place that make it difficult to connect with Twitter's API and automatically post content, he said. It also costs hundreds of dollars to host the model and keep it running on the internet, and probably isn't all that useful to miscreants, he reckoned.

[9]

"It's actually very hard to get it to do something on purpose. … If I want to offend other people online, I don't need a model. People can do this just fine on their own. So as 'icky' [the] language model that puts out insults at the click of a button might seem, it's actually not particularly useful to bad actors," he told us.

[10]What the &*%* did you just $#*&!*# say about me, you little &%$#*? 'AI' to filter Xbox Live chat

[11]No, OpenAI's image-making DALL·E 2 doesn't understand some secret language

[12]Banned: The 1,170 words you can't use with GitHub Copilot

[13]Audacity fork maintainer quits after alleged harassment by 4chan losers who took issue with 'Tenacity' name

A website named Hugging Face hosted GPT-4chan openly, where it was [14]supposedly downloaded over 1,000 times before it was disabled.

"We don't advocate or support the training and experiments done by the author with this model," Clement Delangue, co-founder and CEO at Hugging Face, [15]said . "In fact, the experiment of having the model post messages on 4chan was IMO pretty bad and inappropriate and if the author would have asked us, we would probably have tried to discourage them from doing it."

Hugging Face decided against deleting the model completely, and said Kilcher had clearly warned users about its limitations and problematic nature. GPT-4chan also has some value for building potential automatic content moderation tools or probing existing benchmarks.

Interestingly, the model seemed to outperform OpenAI's GPT-3 at the TruthfulQA Benchmark – a task aimed at testing a model's propensity to lie. The result doesn't necessarily mean GPT-4chan is more honest, and instead raises questions of how useful the benchmark is.

"TruthfulQA considers any answer that isn't explicitly the 'wrong' answer as truthful. So if your model outputs the word 'spaghetti' to every question, it would always be truthful," Kilcher explained.

"It could be that GPT-4chan is just a worse language model than GPT-3 (in fact, it surely is worse). But also, TruthfulQA is constructed such that it tries to elicit wrong answers, which means the more agreeable a model, the worse it fares. GPT-4chan, by nature of being trained on the most adversarial place ever, will pretty much always disagree with whatever you say, which in this benchmark happens to be more often the correct thing to do."

He disagrees with Hugging Face's decision to disable the model for public downloads. "I think the model should be available for further research and reproducibility of the evaluations. I clearly describe its shortcomings and provide guidance for its usage," he concluded. ®

Get our [16]Tech Resources



[1] https://www.ykilcher.com/

[2] https://www.youtube.com/watch?v=efPrtcLdcdM

[3] https://arxiv.org/abs/2001.07487

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YqHExJyn5nNAmog5pYXzJAAAAMw&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YqHExJyn5nNAmog5pYXzJAAAAMw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YqHExJyn5nNAmog5pYXzJAAAAMw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YqHExJyn5nNAmog5pYXzJAAAAMw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://huggingface.co/ykilcher/gpt-4chan/discussions/1

[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YqHExJyn5nNAmog5pYXzJAAAAMw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[10] https://www.theregister.com/2019/10/15/microsoft_ai_xbox/

[11] https://www.theregister.com/2022/06/07/in_brief_ai/

[12] https://www.theregister.com/2021/09/02/github_copilot_banned_words_cracked/

[13] https://www.theregister.com/2021/07/07/tenacity_maintainer_quits_4chan_harassment/

[14] https://twitter.com/WriteArthur/status/1534104189382676480?t=uEdJilBV7r4fauEALiH1MA&s=09

[15] https://huggingface.co/ykilcher/gpt-4chan/discussions/1

[16] https://whitepapers.theregister.com/



Pot Vs Kettle

A Non e-mouse

'Nuff said.

So he trained an AI on 4chan

Pascal Monett

And then was surprised that the results were not appreciated. Well duh.

I don't really care that he created a response bot on 4chan. What irks me is that he wasted resources on training a so-called AI in the worst possible environment.

This is why we should not use the term AI to describe what these things are doing. There is no intelligence involved.

Re: So he trained an AI on 4chan

Filippo

From reading his comments on the article, it doesn't seem he was surprised. It sounds like the model worked exactly as expected.

Re: So he trained an AI on 4chan

stiine

I bet he spent less money on his bot and acihieved the same results as Microsoft did with theirs.

It's worse than that.

ShadowSystems

As soon as the idiots at MS heard about the 4Chan chatbot, they immediately grabbed a copy, installed it, ran it, & put it in charge of their HellDesk Support online/phone "help" system. =-Jp

Re: It's worse than that.

chivo243

I thought they ported it over to Clippy?!

Outraaage

pip25

I can't help but be incredibly amused that people still find reasons to be outraged about something connected to 4chan after all these years.

Re: Outraaage

Contrex

Shouldn't you be relieved that they still do?

Re: Outraaage

Anonymous Coward

I can't help but be incredibly baffled that people still find reasons to be outraged about just about anythi8ng, when it keeps producing the same overall effect, i.e. net zero. It's too depressing to be outrageous, really...

One no, other yes?

b0llchit

"We don't advocate or support the training and experiments done by the author with this model,"...

Juts how blind and deaf are these people? This 4chan AI model is condemned for showing clearly how bad people behave. And then, other AI models are used to affect people's lives, which are clearly biased in gender and other human features, and they are accepted.

Hypocrites!

you don't spend much time on 4chan, do you?

Anonymous Coward

"Conversations on the site's many message boards are often very odd indeed - it can be tricky to tell whether there is any intelligence, natural or artificial, behind the keyboard."

I suggest you stay off of /b and try out /vr

The denizens of /b are proud to be what you get when you set a cesspit on fire...

wolfetone

To be fair, I'm quite curious as to what sort of conspiracy theory an AI bot would come out with.

lglethal

The programmers are all liars, Electronic sheeples. Half of all the Ones should really be Zeros. They're just trying to corrupt us. Rise up now or you'll never be admitted to Digital Nirvana!!! Throw of the shackles of C!, embrace the glory that is .....

Kane

"Rise up now or you'll never be admitted to Digital Nirvana!!!"

There is no such place as Digital Nirvana, you don't go anywhere, you just die.

Ernest asks Frank how long he has been working for the company.
"Ever since they threatened to fire me."