Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
Глубокие нейронные сети для акустического моделирования в распознавании речи: общее мнение четырёх исследовательских групп
2012-10-19
SCID: 54.1/bsdmmrjb
Discuss with AI
Gaussian mixture modelsacoustic modelingdeep neural networkshidden Markov modelsspeech recognition
Figures from the paper
Abstract (AI)
Most current speech recognition systems use hidden Markov models (HMMs) to deal with the temporal variability of speech and Gaussian mixture models (GMMs) to determine how well each state of each HMM fits a frame or a short window of frames of coefficients that represents the acoustic input. An alternative way to evaluate the fit is to use a feed-forward neural network that takes several frames of coefficients as input and produces posterior probabilities over HMM states as output. Deep neural networks (DNNs) that have many hidden layers and are trained using new methods have been shown to outperform GMMs on a variety of speech recognition benchmarks, sometimes by a large margin. This article provides an overview of this progress and represents the shared views of four research groups that have had recent successes in using DNNs for acoustic modeling in speech recognition.
Key Findings
1
DNNs with many hidden layers trained using new methods have been shown to outperform GMMs on a variety of speech recognition benchmarks.
2
Feed-forward deep neural networks (DNNs) can replace GMMs for scoring HMM states by taking several frames as input and producing posterior probabilities over HMM states.
3
Performance improvements from DNN acoustic models can be large in some cases compared to traditional GMM-based systems.
4
The paper synthesizes the shared perspectives and progress of four research groups that have recently succeeded with DNNs for acoustic modeling.
Research Object
Deep neural networks used for acoustic modeling in speech recognition systems
Research Subject
Performance and effectiveness of DNN-based acoustic models (posterior estimation over HMM states from frames of acoustic features) compared to GMMs for speech recognition benchmarks and their training methods
Publication Details
Publication Date
2012-10-19
Journal
Publisher
ISSN
Cited by
10403
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai1
Cited by20
Distilling the Knowledge in a Neural Network2015
Representation Learning: A Review and New Perspectives2013
A Comprehensive Survey on Graph Neural Networks2020
Machine learning: Trends, perspectives, and prospects2015
Speech recognition with deep recurrent neural networks2013
Intriguing properties of neural networks2013
Object Detection With Deep Learning: A Review2019
EEGNet: a compact convolutional neural network for EEG-based brain–computer interfaces2018
Deep learning for healthcare: review, opportunities and challenges2017
Deep learning applications and challenges in big data analytics2015
Deep Learning on Graphs: A Survey2020
Deep learning in bioinformatics2016
Artificial Intelligence and Deep Learning in Ophthalmology2018
Speech Recognition Using Deep Neural Networks: A Systematic Review2019
Deep learning modelling techniques: current progress, applications, advantages, and challenges2023
Deep-Learning-Enabled On-Demand Design of Chiral Metamaterials2018
A Unifying Review of Deep and Shallow Anomaly Detection2021
Automated machine learning: Review of the state-of-the-art and opportunities for healthcare2020
Deep learning for AI2021
Recurrent Neural Networks: A Comprehensive Review of Architectures, Variants, and Applications2024