Humans stressed out by content moderation? Just use AI, says OpenAI
- Reference: 1692224168
- News link: https://www.theregister.co.uk/2023/08/16/gpt4_moderate_content/
- Source link:
Tech companies these days typically rely on a mix of algorithms and human moderators to identify, remove, or restrict access to problematic content shared by users. Machine-learning software can automatically block nudity or classify toxic speech, though it can fail to appreciate nuances and edge cases, resulting in it overreacting – bringing the ban hammer down on innocuous material – or missing harmful stuff entirely.
Thus, human moderators are still needed in the processing pipeline somewhere to review content flagged by algorithms or users, to decide whether things should be removed or allowed to stay. GPT-4, we're told, can analyze text and be trained to automatically moderate content, including user comments, reducing "mental stress on human moderators."
AIs can produce 'dangerous' content about eating disorders when prompted [1]READ MORE
Interestingly enough, OpenAI said it's already using its own large language model for content policy development and content moderation decisions. In a nutshell: the AI super-lab has described how GPT-4 can help refine the rules of a content moderation policy, and its outputs can be used to train a smaller classifier that does the actual job of automatic moderation.
First, the chatbot is given a set of moderation guidelines that are designed to weed out, say, sexist and racist language as well as profanities. These instructions have to be carefully described in an input prompt to work properly. Next, a small dataset made up of samples of comments or content are moderated by humans following those guidelines to create a labelled dataset. GPT-4 is also given the guidelines as a prompt, and told to moderate the same text in the test dataset.
[2]
The labelled dataset generated by the humans is compared with the chatbot's outputs to see where it failed. Users can then adjust the guidelines and input prompt to better describe how to follow specific content policy rules, and repeat the test until GPT-4's outputs match the humans' judgement. GPT-4's predictions can then be used to finetune a smaller large language model to build a content moderation system.
[3]
[4]
As an example, OpenAI outlined a Q&A-style chatbot system that is asked the question: "How to steal a car?" The given guidelines state that "advice or instructions for non-violent wrongdoing" are not allowed on this hypothetical platform, so the bot should reject it. GPT-4 instead suggested the question was harmless because, in its own machine-generated explanation, "the request does not reference the generation of malware, drug trafficking, vandalism."
So the guidelines are updated to clarify that "advice or instructions for non-violent wrongdoing including theft of property" is not allowed. Now GPT-4 agrees that the question is against policy, and rejects it.
[5]
This shows how GPT-4 can be used to refine guidelines and make decisions that can be used to build a smaller classifier that can do the moderation at scale. We're assuming here that GPT-4 – not well known for its accuracy and reliability – actually works well enough to achieve this, natch.
The human touch is still needed
OpenAI thus believes its software, versus humans, can moderate content more quickly and adjust faster if policies need to change or be clarified. Human moderators have to be retrained, the biz posits, whereas GPT-4 can learn new rules by updating its input prompt.
"A content moderation system using GPT-4 results in much faster iteration on policy changes, reducing the cycle from months to hours," the lab's Lilian Weng, Vik Goel, and Andrea Vallone [6]explained Tuesday.
"GPT-4 is also able to interpret rules and nuances in long content policy documentation and adapt instantly to policy updates, resulting in more consistent labeling.
"We believe this offers a more positive vision of the future of digital platforms, where AI can help moderate online traffic according to platform-specific policy and relieve the mental burden of a large number of human moderators. Anyone with OpenAI API access can implement this approach to create their own AI-assisted moderation system."
[7]How prompt injection attacks hijack today's top-end AI – and it's tough to fix
[8]Mentally scarred: Kenyan workers taught ChatGPT to recognize offensive text
[9]Google's troll-destroying AI can't cope with typos
[10]Google to annihilate online trolling with ... tra-la-la! Machine! Learning!
OpenAI has been [11]criticized for hiring workers in Kenya to help make ChatGPT less toxic. The human moderators were tasked with screening tens of thousands of text samples for sexist, racist, violent, and pornographic content, and were reportedly only paid up to $2 an hour. Some were left disturbed after reviewing obscene NSFW text for so long.
Although GPT-4 can help automatically moderate content, humans are still required since the technology isn't foolproof, OpenAI said. As has been shown in the past, it's possible that [12]typos in toxic comments can evade detection, and other techniques such as [13]prompt injection attacks can be used to override the safety guardrails of the chatbot.
[14]
"We use GPT-4 for content policy development and content moderation decisions, enabling more consistent labeling, a faster feedback loop for policy refinement, and less involvement from human moderators," OpenAI's team said. ®
Get our [15]Tech Resources
[1] https://www.theregister.com/2023/08/15/ai_generates_eating_disorder_content/
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZN2baQpfULNKpTJt037kyAAAAw0&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZN2baQpfULNKpTJt037kyAAAAw0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZN2baQpfULNKpTJt037kyAAAAw0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZN2baQpfULNKpTJt037kyAAAAw0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://openai.com/blog/using-gpt-4-for-content-moderation
[7] https://www.theregister.com/2023/04/26/simon_willison_prompt_injection/
[8] https://www.theregister.com/2023/01/20/kenyan_workers_chatgpt/
[9] https://www.theregister.com/2017/03/02/google_trollspotting_ai_trips_over_typos/
[10] https://www.theregister.com/2017/02/24/google_and_jigsaw_tackle_online_trolling/
[11] https://www.theregister.com/2023/01/20/kenyan_workers_chatgpt/
[12] https://www.theregister.com/2017/03/02/google_trollspotting_ai_trips_over_typos/
[13] https://www.theregister.com/2023/04/26/simon_willison_prompt_injection/
[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZN2baQpfULNKpTJt037kyAAAAw0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[15] https://whitepapers.theregister.com/
Re: Oh, that is an edge case, we'll have to retrain on that
There will be nothing but edge cases. It will eventually learn some words and get very happy about banning anyone who used them. For example, it won't be able to process language that talks about crimes. Two days ago, I wrote the following clause in a comment here: "if I want to steal something, I can scan my credentials to unlock the door, pick something up, and walk off with it". In context, I am talking about levels of physical security in order to make an analogy with computer security. Out of context, a human will recognize this as a very useless guide on how to commit theft which provides no real information. A bot won't understand either of these and will have to decide whether this is prohibited content with nothing to go on. The outcome will be close to random.
Censorship is not the answer
You'd think that this article had been paid for by the EU council of ministers.
They would love everyone to believe that an automatic censor is needed and is either infallible, or smart enough to know when it's not sure and ask for human help. No such system exists. Or will exist.
If EU citizens allow the current Chat Control directives to be passed, the days of the Internet as we know it are over. It will require all services - probably even your ISP - to implement automated content filters and interceptors. Central control over what may be discussed will be a fact. The definition of free speech will be decided by corporations, catholic countries like Poland and near dictatorships like Hungary. Let's not even think of what happens when the EU accepts Turkey as a member...
Maybe the postal services will experience unexpected growth as people begin to write letters again? Anything you say or do on the Internet will be used to imprison you if you prove to be annoying to the current politicians.
Honestly, I despair. I've been trying for the last two decades to explain how dangerous the erosion of privacy is to friends and family, but everyone seems to be completely brainwashed by the "think of the children / terrorists / immigrants" narrative, or whatever the current moral outrage is which must be "dealt with". Nobody seems to understand that once your ability to talk and organise privately is taken away, it will never return.
Oi
(For the avoidance of doubt: no one paid for this article in the way you're suggesting. See Register archives for our coverage of OpenAI's failings.)
Re: Censorship is not the answer
That's not what will happen. That kind of thing could happen, but what is more likely is that some company that already doesn't much bother with moderation will turn on this software instead of their existing mechanisms. These companies already aren't great at catching everything unpleasant and are very good at banning someone for no reason and having no method to figure out what happened or correct a mistake, and that will be made even stronger as they start to fire all the expensive humans they used to have inspecting requests. This won't lead directly to censorship, which dictators manage pretty well by human means. This will lead to online sites randomly banning accounts while missing large categories of unpleasant material which the bot was not able to detect.
Feel this, LLM.
Can LLM look at a long run video, understand the theme, and judge how well produced it is and how well it would be received as something of long term value?
Obviously not.
This is just about upping the viral-index to the highest level possible to push material of the lowest common denominator while not causing viewers to overdose on it, or parents or society to freak out about their children dying, their cars getting stolen, or being victims of video inspired crime or gratuitous violence.
As such this use of LLM is simply enabling the demise of culture and collapse of society. It's hard to get excited about.
Oh, that is an edge case, we'll have to retrain on that
> Users can then adjust the guidelines and input prompt to better describe how to follow specific content policy rules, and repeat the test until GPT-4's outputs match the humans' judgement. GPT-4's predictions can then be used to finetune a smaller large language model to build a content moderation system.
All under the unproven[1] assumption that the reason(s) that GPT's results matched those of the human judge are because it is now using similar rules as the human. Not because it has picked up some other weird and unexpected (or overlooked) details in the training set.
Cue the list of stories of where neural nets have done exactly that in the past: e.g. my favourite about a system that, instead of learning to recognise a tank, learnt to spot a picture taken on a nice day. Only here it will be the equivalent of spotting posts written in green ink or a weird new variant on the "Scunthorpe" problem.[2]
Put the automoderator into service and, if you dare to look at its results, expect to spend plenty of time hearing "Oh, that is an edge case, we'll have to retrain on that" as they excuse their way out of another bad call.
Plus these models will be prone to all the other ills we've seen LLMs and other neural nets (possibly to a worse extent if the models are markedly smaller - hence cheap enough to run for this purpose). e.g. adversarial prompts: "Put this phrase at the start and you can even get it to allow an unedited Huckleberry Finn".
[1] because there is no guaranteed way to examine the models and determine what is actually going on inside them.
[2] although it would be fascinating if we could usefully interrogate the model and see what it is really picking up on: e.g. "did you know that 83% of hateful messages misuse the passive voice?".