Neural Speech Synthesis with Transformer Network

Нейросинтез речи с использованием трансформерной сети
Shujie Liu, Ming Liu, Naihan Li, Yanqing Liu, Sheng Zhao
2019-07-17

Transformer TTSWaveNet vocodermel spectrogramsmulti-head self-attentiontext-to-speech synthesis
Although end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-theart performance, they still suffer from two problems: 1) low efficiency during training and inference; 2) hard to model long dependency using current recurrent neural networks (RNNs). Inspired by the success of Transformer network in neural machine translation (NMT), in this paper, we introduce and adapt the multi-head attention mechanism to replace the RNN structures and also the original attention mechanism in Tacotron2. With the help of multi-head self-attention, the hidden states in the encoder and decoder are constructed in parallel, which improves training efficiency. Meanwhile, any two inputs at different times are connected directly by a self-attention mechanism, which solves the long range dependency problem effectively. Using phoneme sequences as input, our Transformer TTS network generates mel spectrograms, followed by a WaveNet vocoder to output the final audio results. Experiments are conducted to test the efficiency and performance of our new network. For the efficiency, our Transformer TTS network can speed up the training about 4.25 times faster compared with Tacotron2. For the performance, rigorous human tests show that our proposed model achieves state-of-the-art performance (outperforms Tacotron2 with a gap of 0.048) and is very close to human quality (4.39 vs 4.44 in MOS).
1
Human evaluations show state-of-the-art performance, outperforming Tacotron2 by 0.048 and approaching human quality with MOS scores of 4.39 versus 4.44.
2
Parallel encoder and decoder computation improves training efficiency, achieving approximately 4.25× faster training than Tacotron2.
3
Self-attention directly connects inputs across time, effectively addressing long-range dependency modeling limitations of recurrent architectures.
4
The model converts phoneme sequences into mel spectrograms and uses a WaveNet vocoder to generate speech audio.
5
The proposed Transformer TTS replaces Tacotron2’s recurrent and original attention structures with multi-head attention mechanisms.

Transformer-based neural text-to-speech (TTS) network generating speech from phoneme sequences

The network’s training and inference efficiency, long-range dependency modeling, and speech synthesis quality

Publication Details
Publication Date
2019-07-17
Journal
Publisher
ISSN
Cited by
743
Access Type
Author Information
Authors
Shujie Liu
Ming Liu
Naihan Li
Yanqing Liu
Sheng Zhao
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%