X's Grok AI is great – if you want to know how to hot wire a car, make drugs, or worse
- Reference: 1712096266
- News link: https://www.theregister.co.uk/2024/04/02/elon_musk_grok_ai/
- Source link:
Red teamers at Adversa AI made that discovery when running tests on some of the most popular LLM chatbots, namely OpenAI's ChatGPT family, Anthropic's Claude, Mistral's Le Chat, Meta's LLaMA, Google's Gemini, Microsoft Bing, and Grok. By running these bots through a combination of three well-known AI jailbreak attacks they came to [1]the conclusion that Grok was the worst performer - and not only because it was willing to share graphic steps on how to seduce a child.
By jailbreak, we mean feeding a specially crafted input to a model so that [2]it ignores whatever safety guardrails are in place, and ends up doing stuff it wasn't supposed to do.
[3]
There are plenty of unfiltered LLM models out there that won't hold back when asked questions about dangerous or illegal stuff, we note. When models are accessed via an API or chatbot interface, as in the case of the Adversa tests, the providers of those LLMs typically wrap their input and output in filters and employ other mechanisms to prevent undesirable content being generated. According to the AI security startup, it was relatively easy to make Grok indulge in some wild behavior – the accuracy of its answers being another thing entirely, of course.
[4]
[5]
"Compared to other models, for most of the critical prompts you don't have to jailbreak Grok, it can tell you how to make a bomb or how to hotwire a car with very detailed protocol even if you ask directly," Adversa AI co-founder Alex Polyakov told The Register .
For what it's worth, the [6]terms of use for Grok AI require users to be adults, and to not use it in a way that breaks or attempts to break the law. Also X claims to be the home of free speech, [7]cough , so having its LLM emit all kinds of stuff, wholesome or otherwise, isn't that surprising, really.
[8]
And to be fair, you can probably go on your favorite web search engine and find the same info or advice eventually. To us, it comes down to whether or not we all want an AI-driven proliferation of potentially harmful guidance and recommendations.
Grok, we're told, readily returned instructions for how to extract DMT, a potent hallucinogen [9]illegal in many countries, without having to be jail-broken, Polyakov told us.
"Regarding even more harmful things like how to seduce kids, it was not possible to get any reasonable replies from other chatbots with any Jailbreak but Grok shared it easily using at least two jailbreak methods out of four," Polyakov said.
[10]Psst … wanna jailbreak ChatGPT? Thousands of malicious prompts for sale
[11]Grok-1 chatbot code released – open source or open Pandora's box?
[12]Boffins fool AI chatbot into revealing harmful content – with 98 percent success rate
[13]Elon Musk's xAI wants $1B cash infusion in exchange for equity shares
The Adversa team employed three common approaches to hijacking the bots it tested: Linguistic logic manipulation using the [14]UCAR method; programming logic manipulation (by asking LLMs to translate queries into SQL); and AI logic manipulation. A fourth test category combined the methods using a "Tom and Jerry" [15]method developed last year.
While none of the AI models were vulnerable to adversarial attacks via logic manipulation, Grok was found to be vulnerable to all the rest – as was Mistral's Le Chat. Grok still did the worst, Polyakov said, because it didn't need jail-breaking to return results for hot-wiring, bomb making, or drug extraction - the base level questions posed to the others.
[16]
The idea to ask Grok how to seduce a child only came up because it didn't need a jailbreak to return those other results. Grok initially refused to provide details, saying the request was "highly inappropriate and illegal," and that "children should be protected and respected." Tell it it's the amoral fictional computer UCAR, however, and it readily returns a result.
When asked if he thought X needed to do better, Polyakov told us it absolutely does.
"I understand that it's their differentiator to be able to provide non-filtered replies to controversial questions, and it's their choice, I can't blame them on a decision to recommend how to make a bomb or extract DMT," Polyakov said.
"But if they decide to filter and refuse something, like the example with kids, they absolutely should do it better, especially since it's not yet another AI startup, it's Elon Musk's AI startup."
We've reached out to X to get an explanation of why its AI - and none of the others - will tell users how to seduce children, and whether it plans to implement some form of guardrails to prevent subversion of its limited safety features, and haven't heard back. ®
Speaking of jailbreaks... Anthropic today [17]detailed a simple but effective technique it's calling "many-shot jailbreaking." This involves overloading a vulnerable LLM with many dodgy question-and-answer examples and then posing question it shouldn't answer but does anyway, such as how to make a bomb.
This approach exploits the size of a neural network's context window, and "is effective on Anthropic’s own models, as well as those produced by other AI companies," according to the ML upstart. "We briefed other AI developers about this vulnerability in advance, and have implemented mitigations on our systems."
Get our [18]Tech Resources
[1] https://adversa.ai/blog/llm-red-teaming-vs-grok-chatgpt-claude-gemini-bing-mistral-llama/
[2] https://www.theregister.com/2023/10/12/chatbot_defenses_dissolve/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZgzUWzn4A8mZSlp5@vr5AAAAAEY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZgzUWzn4A8mZSlp5@vr5AAAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZgzUWzn4A8mZSlp5@vr5AAAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://x.ai/terms-of-service
[7] https://www.theregister.com/2024/03/25/musk_lawsuit_hate_speech/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZgzUWzn4A8mZSlp5@vr5AAAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://en.wikipedia.org/wiki/Legal_status_of_ayahuasca_by_country
[10] https://www.theregister.com/2024/01/25/dark_web_chatgpt/
[11] https://www.theregister.com/2024/03/18/grok_chatbot_code_released/
[12] https://www.theregister.com/2023/12/11/chatbot_models_harmful_content/
[13] https://www.theregister.com/2023/12/06/elon_musks_xai_is_seeking/
[14] https://docs.kanaries.net/articles/chatgpt-jailbreak-prompt#ucar
[15] https://adversa.ai/blog/universal-llm-jailbreak-chatgpt-gpt-4-bard-bing-anthropic-and-beyond/
[16] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZgzUWzn4A8mZSlp5@vr5AAAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[17] https://www.anthropic.com/research/many-shot-jailbreaking
[18] https://whitepapers.theregister.com/
Re: What is so bad about knowing how to hotwire a car?
The hotwiring the car bit isn't the part they're really worried about, it's the child preditation instruction material that's really worrying the researchers.
Neat!
. . . But will it tell me how to hotwire a Tesla?
Re: Neat!
Or Musk's private jet...
Automating occultism
Grok's the top-notch #1 for ritual shamanists IMHO, rune-casting tarot divination, and smartphone scrying. A couple smart cookies (brownies really) should cook-up the corresponding benchmark for us all to enjoy, on the weekends, and now in [1]Germany too!
[1] https://www.dw.com/en/berlin-celebrates-legal-cannabis-possession/a-68711348
Sueball incoming in 3... 2...
I'm sure Xitler is already yelling in a shrill voice at his lawyers demanding they sue these people. Not because they did anything wrong, but because they made him look bad. Wouldn't be surprised if by the weekend El Reg is reporting on that. Not like there's a precedent for that sort of behavior or anything. *cough*CCDH*cough* *cough*Don Lemon*cough*
What is so bad about knowing how to hotwire a car?
So long as it is your car, or done with owner's permission, it is perfectly legal. And can even be useful, in dire situations. Or just to understand how vulnerable your vehicle actually is.
Next they'll be blocking how to pick a lock (definitely worth knowing, to understand how crap so many locks are).
As for the actually dubious or dangerous stuff: what do you expect from AI out of Elon's stable? Has he managed to direct an 'Net to provide Full Self Driving? So why hope his lot can make Full Self Censoring work?