Neural Speech Synthesis with Transformer Network
Нейросинтез речи с использованием трансформерной сети
2019-07-17
SCID: 54.1/drcfytzg
Discuss with AI
Transformer TTSWaveNet vocodermel spectrogramsmulti-head self-attentiontext-to-speech synthesis
Figures from the paper
Abstract (AI)
Although end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-theart performance, they still suffer from two problems: 1) low efficiency during training and inference; 2) hard to model long dependency using current recurrent neural networks (RNNs). Inspired by the success of Transformer network in neural machine translation (NMT), in this paper, we introduce and adapt the multi-head attention mechanism to replace the RNN structures and also the original attention mechanism in Tacotron2. With the help of multi-head self-attention, the hidden states in the encoder and decoder are constructed in parallel, which improves training efficiency. Meanwhile, any two inputs at different times are connected directly by a self-attention mechanism, which solves the long range dependency problem effectively. Using phoneme sequences as input, our Transformer TTS network generates mel spectrograms, followed by a WaveNet vocoder to output the final audio results. Experiments are conducted to test the efficiency and performance of our new network. For the efficiency, our Transformer TTS network can speed up the training about 4.25 times faster compared with Tacotron2. For the performance, rigorous human tests show that our proposed model achieves state-of-the-art performance (outperforms Tacotron2 with a gap of 0.048) and is very close to human quality (4.39 vs 4.44 in MOS).
Key Findings
1
Human evaluations show state-of-the-art performance, outperforming Tacotron2 by 0.048 and approaching human quality with MOS scores of 4.39 versus 4.44.
2
Parallel encoder and decoder computation improves training efficiency, achieving approximately 4.25× faster training than Tacotron2.
3
Self-attention directly connects inputs across time, effectively addressing long-range dependency modeling limitations of recurrent architectures.
4
The model converts phoneme sequences into mel spectrograms and uses a WaveNet vocoder to generate speech audio.
5
The proposed Transformer TTS replaces Tacotron2’s recurrent and original attention structures with multi-head attention mechanisms.
Research Object
Transformer-based neural text-to-speech (TTS) network generating speech from phoneme sequences
Research Subject
The network’s training and inference efficiency, long-range dependency modeling, and speech synthesis quality
Publication Details
Publication Date
2019-07-17
Journal
Publisher
ISSN
Cited by
743
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai5
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation2014
Neural Machine Translation by Jointly Learning to Align and Translate2014
Sequence to Sequence Learning with Neural Networks2014
Self-Attention Generative Adversarial Networks2018