VALL-E AI can mimic a person’s voice from a three-second snippet
- Reference: 1673512209
- News link: https://www.theregister.co.uk/2023/01/12/microsoft_valle_ai/
- Source link:
The technology – called VALL-E and outlined in a 15-page research [1]paper released this month on the arXiv research site – is a significant step forward for Microsoft. TTS is a highly competitive niche that includes other heavyweights such as Google, Amazon, and Meta.
Redmond is already using artificial intelligence for natural language processing (NLP) through its [2]Nuance business – which it bought for $20 billion last year including both speech recognition and TTS technology. And it's aggressively [3]investing in and using technology from startup OpenAI – including its [4]ChatGPT tool – possibly in its Bing search engine and its Office suite of applications.
[5]
A demo of VALL-E can be [6]found on GitHub.
[7]
[8]
In the paper, the researchers argue that while the rise of neural networks and end-to-end modeling has rapidly improved the technologies around speech synthesis, there are still problems with the similarity of the voices used and the lack of natural speaking patterns in TTS products. They aren't the robotic voices of a decade or two ago, but they also don't come off as completely human either.
Caveats
A lot of work is being put into improving this, but there are serious challenges according to the Microsoft eggheads. Some require clean voice data from a recording studio to capture high-quality speech. And they need to rely on relatively small amounts of training data – large-scale speech libraries found on the internet are not clean enough for the work.
For current zero-shot TTS generators – where the software uses samples not included in the training – the work is complex. It can take hours for the system to apply a person's voice to typed text.
"Instead of designing a complex and specific network for this problem, the ultimate solution is to train a model with large and diverse data as much as possible, motivated by success in the field of text synthesis," the researchers wrote, noting that the amount of data being used in text language models in recent years has grown from 16GB of uncompressed text to about a terabyte.
[9]
VALL-E is "the first language model-based TTS framework leveraging large, diverse, and multi-speaker speech data," according to the boffins.
They trained VALL-E with Libri-Light – an open source dataset from Meta that includes 60,000 hours of English speech with more than 7,000 unique speakers. By comparison, other TTS systems are trained using dozens of hours of single-speaker data or hundreds of hours with data from multiple speakers.
VALL-E can keep the acoustic environment of the voice. So if the snippet of voice used as the acoustic prompt in the model is recorded on the telephone, the synthesized spoken text would also sound like it's coming through the phone.
[10]
The capturing of emotion is similar, the researchers claim. If the seconds of recorded voice of the acoustic prompt is emoting anger, then the synthesized speech based on that voice will also display anger.
The result is a TTS model that outperforms others in such areas as natural sounding speech and speaker similarity. Testing also indicates that "the synthesized speech of unseen speakers is as natural as human recordings," they assert.
The researchers noted some issues that need to be resolved – including that some words in the synthesized speech end up missing, are unclear, or are duplicated. There also isn't enough coverage of speakers with accents, and there needs to be greater diversity in speaking styles.
The global TTS market is estimated to grow to tens of billions of dollars by the end of the decade, with both established players and startups driving development of the technology. Microsoft's Nuance business has its TTS product and the software behemoth offers TTS service in Azure. Amazon has Polly, Meta has Meta-TTS, and Google Cloud also offers a service.
All that makes for a crowded space.
[11]Scientists tricked into believing fake abstracts written by ChatGPT were real
[12]AI-generated phishing emails just got much more convincing
[13]Microsoft may be counting out $10 billion to inject into OpenAI
[14]AI conference and NYC's educators ban papers done by ChatGPT
The rapid improvement in the technology raises various ethical and legal issues. A person's voice could be captured and synthesized for use in a wide range of areas – from ads or spam calls to video games or chatbots. They could also be used in deepfakes, with the voice of a politician or celebrity combined with an image to spread disinformation or foment anger.
Patrick Harr, CEO of anti-phishing firm SlashNext, told The Register TTS could also become yet another tool for cybercriminals, who could use it for vishing campaigns – attacks using fraudulent phone calls or voice messages thought to be from a contact the victim knows. It also could be used in more traditional phishing attacks.
"This technology could be extremely dangerous in the wrong hands," Harr said.
The Microsoft researchers noted the risk of synthesized speak that retains the speaker's identity. They said it would be possible to build a detection model to discern whether an audio clip is real or synthesized using VALL-E.
Harr said that within a few years, everyone could have "a unique digital DNA pattern powered by blockchain that can be applied to their voice, content they write, their virtual avatar, etc. This would make it much harder for threat actors to leverage AI for voice impersonation of company executives for example, because those impersonations will lack the 'fingerprint' of the actual executive."
Here's hoping, anyway. ®
Get our [15]Tech Resources
[1] https://arxiv.org/pdf/2301.02111.pdf
[2] https://www.theregister.com/2022/05/12/nuance_healthcare_ai/
[3] https://www.theregister.com/2023/01/10/microsoft_openai_investment_google/
[4] https://www.theregister.com/2023/01/06/chatgpt_cybercriminals_malicious_code/
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y7-oTVeFiq6RTwSi2P@6NAAAAEY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[6] https://valle-demo.github.io/
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y7-oTVeFiq6RTwSi2P@6NAAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y7-oTVeFiq6RTwSi2P@6NAAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y7-oTVeFiq6RTwSi2P@6NAAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_security/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y7-oTVeFiq6RTwSi2P@6NAAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[11] https://www.theregister.com/2023/01/11/scientists_chatgpt_papers/
[12] https://www.theregister.com/2023/01/11/gpt3_phishing_emails/
[13] https://www.theregister.com/2023/01/10/microsoft_openai_investment_google/
[14] https://www.theregister.com/2023/01/06/ai_conference_nyc_ban/
[15] https://whitepapers.theregister.com/
"This technology could be extremely dangerous in the wrong hands,"
This is NEVER gonna happen!
...
no, I'm NOT gonna add a wink! What d'you mean I will! What are you doin to me, here are you takin me!? No, stop, STO
and they've managed to get blockchan into it. we're done for
It'll be very useful for fan made animated pr0n, but probably little else. :-)
but probably little else
How long before the first claim by a politician that "That's not me, that's faked by cyber scoundrels!"?
Sadly, it could be true. If said scoundrels use GPT-3 to produce a speech for TTS input, it might have enough errors in it to sound like a politician. "ChatGPT, produce a speech on how encryption can be backdoored safely..."
So Johnny Scamstain phones me up, I tell him to piss off, he uses that sample to create a copy of my voice and phone my gran saying it's an emergency and I need money quick and the solution our "expert" from MS recommends is the Blockchain?
The blockchain is not going to help my gran.
It seems to me that a legislative framework built around these kinds of AI tooling must include strict liability of the providers for the way their tool is used. If Johnny Scamstain has used this to con my gran, Microsoft are liable. It is the only thing I can think of that will persuade them to take it seriously.
Accents needed
"...There also isn't enough coverage of speakers with accents..."
English is a world language with many variants. Every nation gets restless if some different variant is forced upon them. The same applies to other languages: Spanish and Arabic especially, German and French also. And others.
there are still problems with the similarity of the voices used
yes, those problems have long posed a challenge to scientistis, so great MS are FINALLY working to solve this!
www.youtube.com/watch?v=MT_u9Rurrqg