A study on data augmentation of reverberant speech for robust speech recognition
Исследование увеличения данных реверберационной речи для устойчивого распознавания речи
2017-03-01
SCID: 54.1/2ndg777z
Discuss with AI
DNN-based acoustic modelsLVCSR tasksdata augmentationfar-field ASRmulti-condition trainingpoint-source noisereverberant speechroom impulse responses (RIRs)
Figures from the paper
Abstract (AI)
The environmental robustness of DNN-based acoustic models can be significantly improved by using multi-condition training data. However, as data collection is a costly proposition, simulation of the desired conditions is a frequently adopted strategy. In this paper we detail a data augmentation approach for far-field ASR. We examine the impact of using simulated room impulse responses (RIRs), as real RIRs can be difficult to acquire, and also the effect of adding point-source noises. We find that the performance gap between using simulated and real RIRs can be eliminated when point-source noises are added. Further we show that the trained acoustic models not only perform well in the distant-talking scenario but also provide better results in the close-talking scenario. We evaluate our approach on several LVCSR tasks which can adequately represent both scenarios.
Key Findings
1
Acoustic models trained with the proposed reverberant augmentation perform well for distant-talking (far-field) recognition and also improve close-talking (near-field) results.
2
Adding point-source noises to augmented reverberant speech closes the performance gap between simulated and real RIRs.
3
Multi-condition training data improves environmental robustness of DNN-based acoustic models, and simulation is a viable strategy when real data collection is costly.
4
The proposed augmentation approach was evaluated across several LVCSR tasks representing both close- and far-field scenarios.
5
Using simulated room impulse responses (RIRs) for data augmentation can match the performance of real RIRs when point-source noises are added.
Research Object
Reverberant speech data used for training DNN-based acoustic models for far-field ASR
Research Subject
Effects of data augmentation via simulated room impulse responses and added point-source noises on robustness and recognition performance of acoustic models in far-field (and close-talking) speech recognition
Publication Details
Publication Date
2017-03-01
Journal
Publisher
ISSN
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai2
Cited by4
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing2022
Self-Supervised Speaker Recognition with Loss-Gated Learning2022
End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-Entropy2022
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018