A study on data augmentation of reverberant speech for robust speech recognition

Исследование увеличения данных реверберационной речи для устойчивого распознавания речи
Daniel Povey, Sanjeev Khudanpur, Tom Ko, Vijayaditya Peddinti, Michael L. Seltzer
2017-03-01

DNN-based acoustic modelsLVCSR tasksdata augmentationfar-field ASRmulti-condition trainingpoint-source noisereverberant speechroom impulse responses (RIRs)
The environmental robustness of DNN-based acoustic models can be significantly improved by using multi-condition training data. However, as data collection is a costly proposition, simulation of the desired conditions is a frequently adopted strategy. In this paper we detail a data augmentation approach for far-field ASR. We examine the impact of using simulated room impulse responses (RIRs), as real RIRs can be difficult to acquire, and also the effect of adding point-source noises. We find that the performance gap between using simulated and real RIRs can be eliminated when point-source noises are added. Further we show that the trained acoustic models not only perform well in the distant-talking scenario but also provide better results in the close-talking scenario. We evaluate our approach on several LVCSR tasks which can adequately represent both scenarios.
1
Acoustic models trained with the proposed reverberant augmentation perform well for distant-talking (far-field) recognition and also improve close-talking (near-field) results.
2
Adding point-source noises to augmented reverberant speech closes the performance gap between simulated and real RIRs.
3
Multi-condition training data improves environmental robustness of DNN-based acoustic models, and simulation is a viable strategy when real data collection is costly.
4
The proposed augmentation approach was evaluated across several LVCSR tasks representing both close- and far-field scenarios.
5
Using simulated room impulse responses (RIRs) for data augmentation can match the performance of real RIRs when point-source noises are added.

Reverberant speech data used for training DNN-based acoustic models for far-field ASR

Effects of data augmentation via simulated room impulse responses and added point-source noises on robustness and recognition performance of acoustic models in far-field (and close-talking) speech recognition

Publication Details
Publication Date
2017-03-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Daniel Povey
Sanjeev Khudanpur
Tom Ko
Vijayaditya Peddinti
Michael L. Seltzer
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%