News: 1650525847

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Machine-learning models vulnerable to undetectable backdoors: new claim

(2022/04/21)


Boffins from UC Berkeley, MIT, and the Institute for Advanced Study in the United States have devised techniques to implant undetectable backdoors in machine learning (ML) models.

Their work suggests ML models developed by third parties fundamentally cannot be trusted.

In [1]a paper that's currently being reviewed – "Planting Undetectable Backdoors in Machine Learning Models" – [2]Shafi Goldwasser , [3]Michael Kim , [4]Vinod Vaikuntanathan , and [5]Or Zamir explain how a malicious individual creating a machine learning classifier – an algorithm that classifies data into categories (eg "spam" or "not spam") – can subvert the classifier in a way that's not evident.

[6]

"On the surface, such a backdoored classifier behaves normally, but in reality, the learner maintains a mechanism for changing the classification of any input, with only a slight perturbation," the paper explains. "Importantly, without the appropriate 'backdoor key,' the mechanism is hidden and cannot be detected by any computationally-bounded observer."

[7]

[8]

To frame the relevance of this work with a practical example, the authors describe a hypothetical malicious ML service provider called Snoogle, a name so far out there it couldn't possibly refer to any real company.

Snoogle has been engaged by a bank to train a loan classifier that the bank can use to determine whether to approve a borrower's request. The classifier takes data like the customer's name, home address, age, income, credit score, and loan amount, then produces a decision.

[9]

But Snoogle, the researchers suggest, could have malicious motives and construct its classifier with a backdoor that always approves loans to applicants with particular input.

"Then, Snoogle could illicitly sell a 'profile-cleaning' service that tells a customer how to change a few bits of their profile, eg the least significant bits of the requested loan amount, so as to guarantee approval of the loan from the bank," the paper explains.

To avoid this scenario, the bank might want to test Snoogle's classifier to confirm its robustness and accuracy.

A backdoored classifier behaves normally, but in reality, the learner maintains a mechanism for changing the classification of any input

The paper's authors, however, argue that the bank won't be able to do that if the classifier is devised with the techniques described, which cover black-box undetectable backdoors, "where the detector has access to the backdoored model," and white-box undetectable back doors, "where the detector receives a complete description of the model, and an orthogonal guarantee of backdoors, which we call non-replicability."

The black-box technique outlined relies on coupling a classifier input with a digital signature. It uses a public-key verification process running alongside the classifier to trigger the backdoor when the message-signature pairs get verified.

[10]

"In all, our findings can be seen as decisive negative results towards current forms of accountability in the delegation of learning: under standard cryptographic assumptions, detecting backdoors in classifiers is impossible," the paper states. "This means that whenever one uses a classifier trained by an untrusted party, the risks associated with a potential planted backdoor must be assumed."

This is such an expansive statement that people taking note of the paper on social media have found it hard to believe, even though the paper includes mathematical proofs.

Read the science

Said one individual [11]on Twitter , "This is false in practice. At least for networks with the ReLu based networks. You can put ReLu based neural networks through a (robust) MILP solver which is guaranteed to discover these backdoors."

The Register put this challenge to two of the paper's authors and both dismissed it.

Or Zamir, a postdoctoral researcher at the Institute for Advanced Study and Princeton University, said that's simply wrong.

"Solving MILP is NP-hard (that is, very unlikely to have an efficient solution always) and thus MILP solvers use heuristics that can't always work, but just work sometimes," said Zamir. "We prove that if you could find our backdoor you could break some very well believed cryptographic assumptions."

Michael Kim, a postdoctoral fellow at UC Berkeley, said he doubted the commenter actually read the paper.

"Based on our proofs, there are no practical (existing) or theoretical (future) analyses that will detect these backdoors, unless you break cryptography," he said. "ReLU or otherwise doesn't matter."

"The biggest contribution of our paper is to formalize what we mean by 'undetectable,'" explained Kim. "We make this notion precise through the language of Cryptography and Complexity Theory."

"Undetectability, in this sense, is a property that we *prove* about our constructions. If you believe in the security guaranteed by standard cryptography, eg that the schemes used to perform encryption of files on your computer are secure, then you must also believe in the undetectability of our constructions."

[12]AI drug algorithms can be flipped to invent bioweapons

[13]Study: AI detects backdoor-unlocking DNA samples

[14]Researchers say objects can hide from computer vision by seeking out unusual company that trips correlation bias

[15]Can your AI code be fooled by vandalized images or clever wording? Microsoft open sources a tool to test for that

Asked whether the undetectability of these backdoors will persist as quantum computing matures, both Kim and Zamir expect that's true.

"Our constructions are undetectable even to quantum algorithms (under the current cryptographic beliefs/state of affairs)," said Kim. "Specifically, they can be instantiated under the LWE problem (Learning with Errors) which is the basis of most post-quantum cryptography."

"Our assumptions are lattice-based and are believed to be post-quantum secure," said Zamir.

Assuming these assumptions survive the peer review process, the researcher's work suggests third-party services that create ML models will need to come up with a way to guarantee that their work can be trusted – something the open source software supply chain has not solved.

"What we show is that blind trust of services is very dangerous," said Kim. "The way to make these services trustworthy lies in the field of Delegation of Computation, specifically delegation of learning. Shafi [Goldwasser, the director of the Simons Institute for the Theory of Computing in Berkeley,] is one of the pioneers of this area, which studies how a weak client can delegate computational tasks to an untrusted but powerful service provider."

In other words, the formal undetectability of these backdoor techniques does not preclude adjusting the ML model creation process to compensate.

"The client and service provider engage in an interaction that requires the provider to prove that they performed the computation correctly," explained Kim. "Our work motivates this formal study even more, tailored to the context of learning (which Shafi has [16]initiated )."

Zamir concurred. "The main point is that you'd not be able to use a network you receive as-is," he said.

One potential mitigation described in the paper, Zamir said, is immunization: doing something to the classifier after you receive it to try to neutralize backdoors. Another, he said, is to require a full transcript of the learning procedure and proof the process was done as documented, which isn't ideal for intellectual property protection or efficiency.

Goldwasser advised caution, and noted that she doesn't expect other forms of machine learning, like unsupervised learning, will end up being better from a security standpoint.

"Be very, very careful," she said. "Get your models verified and hopefully be able to have white box access to them." ®

Get our [17]Tech Resources



[1] https://arxiv.org/abs/2204.06974

[2] https://simons.berkeley.edu/people/shafi-goldwasser

[3] https://cs.stanford.edu/~mpkim/

[4] https://www.csail.mit.edu/person/vinod-vaikuntanathan

[5] https://www.ias.edu/scholars/or-zamir

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YmErVV@L9ZRfD4p3BrOuhgAAAQo&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YmErVV@L9ZRfD4p3BrOuhgAAAQo&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YmErVV@L9ZRfD4p3BrOuhgAAAQo&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YmErVV@L9ZRfD4p3BrOuhgAAAQo&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YmErVV@L9ZRfD4p3BrOuhgAAAQo&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[11] https://twitter.com/Anno0770/status/1516713461211963393?s=20&t=uLNWi-kzkvC16Huxdcr5_w

[12] https://www.theregister.com/2022/03/18/ai_weapons_learning/

[13] https://www.theregister.com/2022/02/25/dna-security-healthcare/

[14] https://www.theregister.com/2021/05/07/computer_vision_correlation_bias/

[15] https://www.theregister.com/2021/05/05/microsoft_ai_security/

[16] https://drops.dagstuhl.de/opus/volltexte/2021/13580/

[17] https://whitepapers.theregister.com/



"Said one individual on Twitter"

Pascal Monett

Yes, because Twitter is a vast reference of people who only post about things they are experts in.

Re: "Said one individual on Twitter"

b0llchit

It is the new collective volume of (140 character) knowledge. That should beat any (online) encyclopedia by miles because the attention span of the reader cannot read beyond five words anyway. Therefore, twitter must be right.

IT is a Brave New NEUKlearer HyperRadioProACTive World for Advanced IntelAIgents ‽ .

amanfromMars 1

the authors describe a hypothetical malicious ML service provider called Snoogle, a name so far out there it couldn't possibly refer to any real company.

Crikey, you gotta get out more, Thomas Claburn in San Francisco, if you believe that a strange name/company name bears no possible relation to reality ...... or are you trying to tell us, without actually directly telling us, that the true nature of reality is fundamentally strange and far out there.

Now/Then you're making more than just some sense and things can be moved along at an elevated pace in what is indeed a most novel and noble space with no old places in which to hide if up to no good.

:-) Of course, you could also be having some targeted fun at Google's broad shouldered expense whose business model and applications of the results of their algorithmic search engines put them directly in the cross hairs of snipers for any number of earlier established elite executive office operations as the likes of a Google becomes ever more powerful and leading in competition and opposition to such as were formerly thought almightily unchallenged and unchallengeable.

Those halcyon days for those leaders of the past are long gone though and they aint coming back for those key holding players ......

And no matter what is being done and no matter where everything may end up, will there always be a that and/or those way out ahead of the game taking every advantage of that which is being developed and trialed/trailed, remotely virtually mentoring and monitoring live operational plays.

It is just the way everything is ... and is now starting to be revealed to you ...... for both either your terror or delight.

Capiche, Amigos/Amigas? Do you need more evidence?

Brewster's Angle Grinder

"Michael Kim, a postdoctoral fellow at UC Berkeley, said he doubted the commenter actually read the paper."

It's a 50-page read. I haven't got time to skim it now.

So it's on the stack to be never read.

There are no manifestos like cannon and musketry.
-- The Duke of Wellington