Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement
Устойчивое распознавание диктора на основе одно- и многоканального улучшения речи
2020-01-01
SCID: 54.1/7rzu2bnb
Discuss with AI
GFCCsMVDR beamformerspeaker recognitionspeech enhancementx-vectors
Figures from the paper
Abstract (AI)
Deep neural network (DNN) embeddings for speaker recognition have recently attracted much attention. Compared to i-vectors, they are more robust to noise and room reverberation as DNNs leverage large-scale training. This article addresses the question of whether speech enhancement approaches are still useful when DNN embeddings are used for speaker recognition. We investigate single- and multi-channel speech enhancement for text-independent speaker verification based on x-vectors in conditions where strong diffuse noise and reverberation are both present. Single-channel (monaural) speech enhancement is based on complex spectral mapping and is applied to individual microphones. We use masking-based minimum variance distortion-less response (MVDR) beamformer and its rank-1 approximation for multi-channel speech enhancement. We propose a novel method of deriving time-frequency masks from the estimated complex spectrogram. In addition, we investigate gammatone frequency cepstral coefficients (GFCCs) as robust speaker features. Systematic evaluations and comparisons on the NIST SRE 2010 retransmitted corpus show that both monaural and multi-channel speech enhancement significantly outperform x-vector's performance, and our covariance matrix estimate is effective for the MVDR beamformer.
Key Findings
1
Complex spectral mapping enables single-channel enhancement applied independently to microphones for robust speaker recognition.
2
Masking-based MVDR beamforming and its rank-1 approximation provide multi-channel speech enhancement, supported by a novel mask derivation method from estimated complex spectrograms.
3
Systematic NIST SRE 2010 retransmitted-corpus evaluations show that both monaural and multi-channel enhancement significantly outperform unenhanced x-vector systems.
4
The proposed covariance matrix estimate is effective for MVDR beamforming, and GFCCs are investigated as robust speaker features.
5
The study evaluates whether speech enhancement improves x-vector speaker verification under strong diffuse noise and reverberation.
Research Object
single- and multi-channel speech enhancement for text-independent speaker verification based on x-vectors under strong diffuse noise and reverberation
Research Subject
robustness and verification performance of x-vector speaker recognition, including the effects of monaural and multi-channel enhancement, MVDR beamforming, and GFCC features under noisy reverberant conditions
Publication Details
Publication Date
2020-01-01
Journal
Publisher
ISSN
Access Type
Author Information
Download PDF
Subscribe to digest