News: 1710318846

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Researchers jimmy OpenAI's and Google's closed models

(2024/03/13)


Boffins have managed to pry open closed AI services from OpenAI and Google with an attack that recovers an otherwise hidden portion of transformer models.

The attack partially illuminates a particular type of so-called "black box" model, revealing the embedding projection layer of a transformer model through API queries. The cost to do so ranges from a few dollars to several thousand, depending upon the size of the model being attacked and the number of queries.

No less than 13 computer scientists from Google DeepMind, ETH Zurich, University of Washington, OpenAI, and McGill University have penned [1]a paper describing the attack, which builds upon a model extraction attack technique [2]proposed in 2016.

[3]

"For under $20 USD, our attack extracts the entire projection matrix of OpenAI's ada and babbage language models," the researchers state in their paper. "We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the gpt-3.5-turbo model, and estimate it would cost under $2,000 in queries to recover the entire projection matrix."

[4]

[5]

The researchers have disclosed their findings to OpenAI and Google, both of which are said to have implemented defenses to mitigate the attack. They chose not to publish the size of two OpenAI gpt-3.5-turbo models, which are still in use. The ada and babbage models are both deprecated, so disclosing their respective sizes was deemed harmless.

While the attack does not completely expose a model, the researchers say that it can reveal the model's final [6]weight matrix – or its width, which is often related to the parameter count – and provides information about the model's capabilities that could inform further probing. They explain that being able to obtain any parameters from a production model is surprising and undesirable, because the attack technique may be extensible to recover even more information.

[7]AI models show racial bias based on written dialect, researchers find

[8]Cloudflare wants to put a firewall in front of your LLM

[9]Stack Overflow to charge LLM developers for access to its coding content

[10]BEAST AI needs just a minute of GPU time to make an LLM fly off the rails

"If you have the weights, then you just have the full model," explained Edouard Harris, CTO at Gladstone AI, in an email to The Register . "What Google [et al.] did was reconstruct some parameters of the full model by querying it, like a user would. They were showing that you can reconstruct important aspects of the model without having access to the weights at all."

Access to enough information about a proprietary model might allow someone to replicate it – a scenario that Gladstone AI considered in [11]a report commissioned by the US Department of State titled "Defense in Depth: An Action Plan to Increase the Safety and Security of Advanced AI".

[12]

The report, [13]released yesterday , provides analysis and recommendations for how the government should harness AI and guard against the ways in which it poses a potential threat to national security.

One of the recommendations of the report is "that the US government urgently explore approaches to restrict the open-access release or sale of advanced AI models above key thresholds of capability or total training compute." That includes "[enacting] adequate security measures to protect critical IP including model weights."

Asked about the Gladstone report's recommendations in light of Google's findings, Harris relied, "Basically, in order to execute attacks like these, you need – at least for now – to execute queries in patterns that may be detectable by the company that's serving the model, which is OpenAI in the case of GPT-4. We recommend tracking high level usage patterns, which should be done in a privacy-preserving way, in order to identify attempts to reconstruct model parameters using these approaches."

[14]

"Of course this kind of first-pass defense might become impractical as well, and we may need to develop more sophisticated countermeasures (e.g., slightly randomizing which models serve which responses at any given time, or other approaches). We don't get into that level of detail in the plan itself however." ®

Get our [15]Tech Resources



[1] https://arxiv.org/abs/2403.06634

[2] https://arxiv.org/abs/1609.02943

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZfGHSV7pPAZMXQUlFYWwtwAAAc8&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZfGHSV7pPAZMXQUlFYWwtwAAAc8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZfGHSV7pPAZMXQUlFYWwtwAAAc8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://www.sciencedirect.com/topics/mathematics/weight-matrix

[7] https://www.theregister.com/2024/03/11/ai_models_exhibit_racism_based/

[8] https://www.theregister.com/2024/03/05/cloudflare_firewall_ai/

[9] https://www.theregister.com/2024/03/01/stack_overflow_launches_api_to/

[10] https://www.theregister.com/2024/02/28/beast_llm_adversarial_prompt_injection_attack/

[11] https://www.gladstone.ai/action-plan#action-plan-overview

[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZfGHSV7pPAZMXQUlFYWwtwAAAc8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[13] https://www.prnewswire.com/news-releases/gladstoneai-announces-the-first-ever-ai-action-plan-for-united-states-national-security-commissioned-by-the-us-state-department-302084845.html

[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZfGHSV7pPAZMXQUlFYWwtwAAAc8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[15] https://whitepapers.theregister.com/



My, what drama

Pascal Monett

All that for a statistical analysis machine that invents stuff on the fly, distorts the truth and cannot give all the relevant data properly.

Keep your weights. It's the concept that is flawed.

Re: My, what drama

abend0c4

It's certainly telling that their primary concern is "security measures to protect critical IP" rather than the consequences of its output.

Isn't it ironic

cyberdemon

That a company calling itself "Open" AI, should be so concerned about the black box around their disruptive-yet-useless product

In the words of Wanda Maximoff...

theOtherJT

So. When openAI uses totally public endpoints to collect data that trains the model it's fair use, when security researchers use public endpoints to disclose openAIs inner workings it's a flaw that has to be patched immediately?

...that doesn't seem fair.

If you have the weights, you have the model.

I ain't Spartacus

Surely to say that having the weights gives you the model is incorrect. Unless you also have identical training data.

Otherwise you have the weightings for one set of data, which are going to be somewhere between subtly and totally different if you use different training data with your copy of the model.

After all, the models are still mining their training data for their outputs. We know this because people are using them as glorified search engines - and researchers have maanged to get them to spit out whole sections of copyright text from their training models verbatim. So you'd need at least a similar dataset to make the weightings meaningful.

The idea is to die young as late as possible.
-- Ashley Montague