Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement

Устойчивое распознавание диктора на основе одно- и многоканального улучшения речи
Hassan Taherian, Zhong-Qiu Wang, Jorge Chang, DeLiang Wang
2020-01-01

GFCCsMVDR beamformerspeaker recognitionspeech enhancementx-vectors
Deep neural network (DNN) embeddings for speaker recognition have recently attracted much attention. Compared to i-vectors, they are more robust to noise and room reverberation as DNNs leverage large-scale training. This article addresses the question of whether speech enhancement approaches are still useful when DNN embeddings are used for speaker recognition. We investigate single- and multi-channel speech enhancement for text-independent speaker verification based on x-vectors in conditions where strong diffuse noise and reverberation are both present. Single-channel (monaural) speech enhancement is based on complex spectral mapping and is applied to individual microphones. We use masking-based minimum variance distortion-less response (MVDR) beamformer and its rank-1 approximation for multi-channel speech enhancement. We propose a novel method of deriving time-frequency masks from the estimated complex spectrogram. In addition, we investigate gammatone frequency cepstral coefficients (GFCCs) as robust speaker features. Systematic evaluations and comparisons on the NIST SRE 2010 retransmitted corpus show that both monaural and multi-channel speech enhancement significantly outperform x-vector's performance, and our covariance matrix estimate is effective for the MVDR beamformer.
1
Complex spectral mapping enables single-channel enhancement applied independently to microphones for robust speaker recognition.
2
Masking-based MVDR beamforming and its rank-1 approximation provide multi-channel speech enhancement, supported by a novel mask derivation method from estimated complex spectrograms.
3
Systematic NIST SRE 2010 retransmitted-corpus evaluations show that both monaural and multi-channel enhancement significantly outperform unenhanced x-vector systems.
4
The proposed covariance matrix estimate is effective for MVDR beamforming, and GFCCs are investigated as robust speaker features.
5
The study evaluates whether speech enhancement improves x-vector speaker verification under strong diffuse noise and reverberation.

single- and multi-channel speech enhancement for text-independent speaker verification based on x-vectors under strong diffuse noise and reverberation

robustness and verification performance of x-vector speaker recognition, including the effects of monaural and multi-channel enhancement, MVDR beamforming, and GFCC features under noisy reverberant conditions

Publication Details
Publication Date
2020-01-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Hassan Taherian
Zhong-Qiu Wang
Jorge Chang
DeLiang Wang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%