News Articles

News Articles


Baidu can clone your voice after hearing just a minute of audio
by Edd Gent.

NEURAL networks can now mimic someone’s voice, and all they need for the feat is less than a minute’s worth of their speech.

Researchers at China’s search engine giant Baidu say the technology could create digital duplicate voices for people who have lost the ability to talk. It could also be used to personalise digital assistants, video game characters or automatic speech translation services.

“A mum could easily configure an audio-book reader with her own voice to read bedtime stories for her kids,” says Sercan Arik at Baidu Research, who led the work.

Voice cloning technology has improved rapidly in recent years. Adobe’s VoCo, released in 2016, could mimic someone’s voice using 20 minutes of audio. Last year, Canadian start-up Lyrebird launched a service letting anyone create a digital copy of their voice based on 1 minute of audio.

Baidu’s research builds on its text-to-speech synthesis system Deep Voice, which was trained on more than 800 hours of audio from 2400 speakers. It builds a model of human speech learning what sounds go with what text and also picks up the idiosyncrasies of each speaker it was trained on.

Now the software is able to synthesise a copy of a voice solely based on hearing snatches of the original. The best version needed 100 snippets, each no more than 5 seconds long, the Baidu team says. But one trained on just 10 snippets performed well enough to dupe a voice recognition system more than 95 per cent of the time, and human evaluators gave it 3.16 out of 4 for mimickry (arxiv.org/abs/1802.06006).

“Digital assistants and banks’ telephone services could be vulnerable to synthesised voices”

The team also tried a secondary method that trained a separate model on just the voice to be mimicked. This was less accurate, but Arik says it is also more efficient so could potentially run on a smartphone.

The output is still not totally indistinguishable from the human voice, says Arik, “but it does show a very fundamental breakthrough in that direction”.

Even the best synthesised voices contain telltale digital signals that are easily detected by advanced voice profiling algorithms, says Rita Singh, a voice forensic science expert at Carnegie Mellon University in Pennsylvania.

However, most voice authentication systems – used to secure everything from banking services to smartphones – can be fooled because they rely instead on picking up broad statistical features, she says.

In 2014, Nitesh Saxena, a security researcher at the University of Alabama at Birmingham, showed that a freely available voice morphing tool could trick voice authentication systems 80 to 90 per cent of the time. Unpublished research shows that leading digital assistants and even a major bank’s telephone service remain vulnerable, he says.

But while biometric systems can be improved, our own ability to detect fakes can’t. This raises the spectre of voice synthesis systems aping someone’s voice to commit fraud or sparking fake news by doctoring a politician’s speech.

“Humans will, over time, become even more vulnerable to such attacks,” says Saxena.

Combining that with approaches like the DeepFake algorithm recently used to transplant celebrities’ faces into porn videos could supercharge the problem, Singh says.

“Now the default status is that if there’s any video that sounds too bad or too good to be true, it’s probably a fake,” she adds.

This article appeared in print under the headline “AI hears snippets of you, then clones your voice”

shape shape