Tacotron: Towards End-to-End Speech Synthesis
Tacotron: к энд-ту-энд синтезу речи
2017-08-16
SCID: 54.1/u3c9zfmz
Discuss with AI
Tacotronend-to-end text-to-speechframe-level speech synthesismean opinion scoresequence-to-sequence TTS
Figures from the paper
Abstract (AI)
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design choices.In this paper, we present Tacotron, an end-to-end generative text-to-speech model that synthesizes speech directly from characters.Given pairs, the model can be trained completely from scratch with random initialization.We present several key techniques to make the sequence-tosequence framework perform well for this challenging task.Tacotron achieves a 3.82 subjective 5-scale mean opinion score on US English, outperforming a production parametric system in terms of naturalness.In addition, since Tacotron generates speech at the frame level, it's substantially faster than sample-level autoregressive methods.
Key Findings
1
Authors introduce key sequence-to-sequence techniques that enable high-performance TTS for this challenging task.
2
Because Tacotron generates speech at the frame level, it is substantially faster than sample-level autoregressive methods.
3
Tacotron achieves a 3.82 mean opinion score (5-point scale) on US English, outperforming a production parametric system in naturalness.
4
Tacotron is an end-to-end generative text-to-speech model that synthesizes speech directly from characters without separate frontend or acoustic modules.
5
The model can be trained from scratch with random initialization using paired <text, audio> data.
Research Object
Tacotron end-to-end generative text-to-speech model that synthesizes speech directly from characters
Research Subject
Performance and characteristics of the model for text-to-speech: ability to be trained from <text,audio> pairs from scratch, sequence-to-sequence techniques for this task, subjective naturalness (MOS), and generation speed at the frame level compared to sample-level autoregressive and parametric systems
Publication Details
Publication Date
2017-08-16
Journal
Publisher
ISSN
Cited by
1750
Access Type
Author Information
Download PDF
Subscribe to digest