Tacotron: Towards End-to-End Speech Synthesis

Tacotron: к энд-ту-энд синтезу речи
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Khai Le, Yannis Agiomyrgiannakis, Rob Clark, Rif A. Saurous
2017-08-16

Tacotronend-to-end text-to-speechframe-level speech synthesismean opinion scoresequence-to-sequence TTS
A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design choices.In this paper, we present Tacotron, an end-to-end generative text-to-speech model that synthesizes speech directly from characters.Given pairs, the model can be trained completely from scratch with random initialization.We present several key techniques to make the sequence-tosequence framework perform well for this challenging task.Tacotron achieves a 3.82 subjective 5-scale mean opinion score on US English, outperforming a production parametric system in terms of naturalness.In addition, since Tacotron generates speech at the frame level, it's substantially faster than sample-level autoregressive methods.
1
Authors introduce key sequence-to-sequence techniques that enable high-performance TTS for this challenging task.
2
Because Tacotron generates speech at the frame level, it is substantially faster than sample-level autoregressive methods.
3
Tacotron achieves a 3.82 mean opinion score (5-point scale) on US English, outperforming a production parametric system in naturalness.
4
Tacotron is an end-to-end generative text-to-speech model that synthesizes speech directly from characters without separate frontend or acoustic modules.
5
The model can be trained from scratch with random initialization using paired <text, audio> data.

Tacotron end-to-end generative text-to-speech model that synthesizes speech directly from characters

Performance and characteristics of the model for text-to-speech: ability to be trained from <text,audio> pairs from scratch, sequence-to-sequence techniques for this task, subjective naturalness (MOS), and generation speed at the frame level compared to sample-level autoregressive and parametric systems

Publication Details
Publication Date
2017-08-16
Journal
Publisher
ISSN
Cited by
1750
Access Type
Author Information
Authors
Yuxuan Wang
RJ Skerry-Ryan
Daisy Stanton
Yonghui Wu
Ron J. Weiss
Navdeep Jaitly
Zongheng Yang
Ying Xiao
Zhifeng Chen
Samy Bengio
Quoc Khai Le
Yannis Agiomyrgiannakis
Rob Clark
Rif A. Saurous
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%