Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups

Глубокие нейронные сети для акустического моделирования в распознавании речи: общее мнение четырёх исследовательских групп
Andrew Senior, George E. Dahl, Geoffrey E. Hinton, Vincent Vanhoucke, Tara N. Sainath, Dong Yu, Abdelrahman Mohamed, Li Deng, Navdeep Jaitly, Patrick Nguyen, Brian Kingsbury
2012-10-19

Gaussian mixture modelsacoustic modelingdeep neural networkshidden Markov modelsspeech recognition
Most current speech recognition systems use hidden Markov models (HMMs) to deal with the temporal variability of speech and Gaussian mixture models (GMMs) to determine how well each state of each HMM fits a frame or a short window of frames of coefficients that represents the acoustic input. An alternative way to evaluate the fit is to use a feed-forward neural network that takes several frames of coefficients as input and produces posterior probabilities over HMM states as output. Deep neural networks (DNNs) that have many hidden layers and are trained using new methods have been shown to outperform GMMs on a variety of speech recognition benchmarks, sometimes by a large margin. This article provides an overview of this progress and represents the shared views of four research groups that have had recent successes in using DNNs for acoustic modeling in speech recognition.
1
DNNs with many hidden layers trained using new methods have been shown to outperform GMMs on a variety of speech recognition benchmarks.
2
Feed-forward deep neural networks (DNNs) can replace GMMs for scoring HMM states by taking several frames as input and producing posterior probabilities over HMM states.
3
Performance improvements from DNN acoustic models can be large in some cases compared to traditional GMM-based systems.
4
The paper synthesizes the shared perspectives and progress of four research groups that have recently succeeded with DNNs for acoustic modeling.

Deep neural networks used for acoustic modeling in speech recognition systems

Performance and effectiveness of DNN-based acoustic models (posterior estimation over HMM states from frames of acoustic features) compared to GMMs for speech recognition benchmarks and their training methods

Publication Details
Publication Date
2012-10-19
Journal
Publisher
ISSN
Cited by
10403
Access Type
Author Information
Authors
Andrew Senior
George E. Dahl
Geoffrey E. Hinton
Vincent Vanhoucke
Tara N. Sainath
Dong Yu
Abdelrahman Mohamed
Li Deng
Navdeep Jaitly
Patrick Nguyen
Brian Kingsbury
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%