Adversarial audio samples generated by AI can trick authentication systems
- Reference: 1687933814
- News link: https://www.theregister.co.uk/2023/06/28/adversarial_ai_audio_fools_authentication/
- Source link:
Voice authentication is commonly used by call centers, banks, government agencies. But AI means attacking such systems has become easier, with researchers claiming a 99 percent success rate for subverting such security.
So a pair of computer scientists from the University of Waterloo in Canada have developed a [1]technique to trick these systems too. A research paper [2]published in the proceedings of the 44th IEEE Symposium on Security and Privacy describes fudging AI-generated speech recordings to create "adversarial" samples that were highly effective.
[3]
Voice authentication relies on the fact that everyone's voice is unique, thanks to physical characteristics like the size and shape of the vocal tract and larynx, and social factors like accent.
[4]
[5]
Voice authentication systems capture those nuances in voiceprints. Although AI-generated audio can fairly realistically mimic people's voices, AI algorithms have their own distinctive artifacts that analysts can spot artificially created voices. The technique developed by the researchers tries to strip these features away, while preserving the overall sound.
"The idea is to 'engrave' the user's voiceprint into the spoofed sample," researchers Andre Kassis and Urs Hengartner wrote in their paper. "Our adversarial engine attempts to remove machine artifacts that are predominant in these samples."
[6]Google warns its own employees: Do not use code generated by Bard
[7]Amazon Ring, Alexa accused of every nightmare IoT security fail you can imagine
[8]UK's GDPR replacement could wipe out oversight of live facial recognition
[9]Vietnam to require registration of social media, even on global platforms
The researchers trained their system on samples of 107 speakers' utterances to get a better idea of what makes speech sound human. To test their algorithm, they crafted multiple adversarial samples to fool authentication systems – with a 72 percent success rate. Against some fairly weak systems, they achieved a 99 percent success rate after six attempts.
This doesn't mean voice authentication software is defunct just yet, though. Against Amazon Connect – software provided to cloud contact centers – they achieved only ten percent success in a four-second attack, and 40 percent in less than 30 seconds. And authentication software is improving too.
[10]
Miscreants hoping to carry out these types of attacks need to have access to their target's voice, and be sufficiently tech-savvy enough to generate their own adversarial audio samples if they're trying to crack a more secure system. Although the barrier is high, the researchers warned companies developing voice authentication software to keep working.
"The success rates of our attacks are concerning," they wrote, "primarily due to them being attained in the black-box setting and under the assumptions of realistic threat models." The findings "highlight the severe pitfalls of voice authentication systems and stress the need for more reliable mechanisms." ®
Get our [11]Tech Resources
[1] https://uwaterloo.ca/news/media/how-secure-are-voice-authentication-systems-really
[2] https://www.computer.org/csdl/proceedings-article/sp/2023/933600a951/1NrbYmtLXB6
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZJwEwJvfLSyJDQIXBpKk3wAAAgY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJwEwJvfLSyJDQIXBpKk3wAAAgY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJwEwJvfLSyJDQIXBpKk3wAAAgY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2023/06/19/even_google_warns_its_own/
[7] https://www.theregister.com/2023/06/01/ftc_alexa_ring_amazon_settlement/
[8] https://www.theregister.com/2023/05/19/dpib_2_surveillance_oversight/
[9] https://www.theregister.com/2023/05/09/vietnam_social_media_registration/
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJwEwJvfLSyJDQIXBpKk3wAAAgY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[11] https://whitepapers.theregister.com/
Re: "need to have access to their target's voice"
While I cannot say from first-hand knowledge, I suspect that the bigger the sample you have, the more accurately the AI system can reproduce it. A few seconds of sampling is unlikely to be enough since most (if not all) of us have idiosyncrasies in our speech patterns that tend to only come out with certain words (in my case, traces of a Canadian accent from when I lived there as a child, even though I came back to Blighty 50-odd years ago).
Re: "need to have access to their target's voice"
Sure, but if you were wanting to target say a politician, there's going to be plenty of material out there to sample.
Re: "need to have access to their target's voice"
>>just how well recorded does the target's voice need to be
Not particularly well. Phone calls have very restricted bandwidth by design (though, recently, you can get 'high resolution' calling if both ends suport it)
If you are serious in this field, you participate in the biannual Automatic Speaker Verification and Spoofing Countermeasures Challenge (ASV spoof):
'https://www.asvspoof.org/
I am very wary about claims of breaking systems in your own lab on your own terms. I want to see it against real systems in a challenge.
My Voice is my Password...
... [1]Verify Me
[1] https://www.youtube.com/watch?v=-zVgWpVXb64
Re: My Voice is my Password...
That movie is the reason I've never trusted voice recognition....
Confused...
Why didn't the system authenticate the tape played at double speed?
On the plus side
According to Noah Vosen it is far better to have your voice recorded than lose a hand or an eye.
Automation giveth, and automation taketh away
... my 2 cents
"need to have access to their target's voice"
With all the audio data available from all sorts of sources on the Internet, that doesn't seem to be much of a barrier if you're seeking to spoof the voice of anyone who is known. Celebrities, politicians, major CEOs, all of them have their voice on publicly-available sources somewhere.
Now, if you're targetting someone for specific reasons that is not a social media aficionado, it makes things a lot more complicated, especially if you do not know the person socially. You're going to have to find a way to meet the person, get the person talking and put your mobile phone down to record the conversation. That will mean cleaning up the recording afterwards, which is never an easy task.
So, the basic question really is just how well recorded does the target's voice need to be ? Will a few dozen seconds in the street suffice, or do you need a few minutes of sound booth recording ?