X-Vectors: Robust DNN Embeddings for Speaker Recognition
X-векторы: устойчивые DNN-эмбеддинги для распознавания говорящего
2018-04-01
SCID: 54.1/kesn7yqz
Discuss with AI
DNN embeddingsPLDA classifierdata augmentationspeaker recognitionx-vectors
Figures from the paper
Abstract (AI)
In this paper, we use data augmentation to improve performance of deep neural network (DNN) embeddings for speaker recognition. The DNN, which is trained to discriminate between speakers, maps variable-length utterances to fixed-dimensional embeddings that we call x-vectors. Prior studies have found that embeddings leverage large-scale training datasets better than i-vectors. However, it can be challenging to collect substantial quantities of labeled data for training. We use data augmentation, consisting of added noise and reverberation, as an inexpensive method to multiply the amount of training data and improve robustness. The x-vectors are compared with i-vector baselines on Speakers in the Wild and NIST SRE 2016 Cantonese. We find that while augmentation is beneficial in the PLDA classifier, it is not helpful in the i-vector extractor. However, the x-vector DNN effectively exploits data augmentation, due to its supervised training. As a result, the x-vectors achieve superior performance on the evaluation datasets.
Key Findings
1
Because of supervised training, x-vector DNNs effectively exploit augmentation and outperform i-vector baselines on Speakers in the Wild and NIST SRE 2016 Cantonese.
2
Data augmentation benefits the PLDA classifier but does not improve the i-vector extractor in the reported experiments.
3
Noise and reverberation augmentation efficiently expand training data and improve the robustness of supervised DNN speaker embeddings.
4
The paper introduces x-vectors, fixed-dimensional speaker embeddings produced by a DNN that maps variable-length utterances while discriminating among speakers.
5
The results show that x-vectors can leverage augmented training data more effectively than i-vectors for speaker recognition.
Research Object
DNN-based x-vector speaker embeddings for speaker recognition
Research Subject
Robustness and recognition performance under noise-and-reverberation data augmentation, compared with i-vector baselines
Publication Details
Publication Date
2018-04-01
Journal
Publisher
ISSN
Cited by
2714
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai5
Cited by9
Voxceleb: Large-scale speaker verification in the wild2019
Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 20202021
Self-Supervised Speaker Recognition with Loss-Gated Learning2022
End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-Entropy2022
Analysis of Length Normalization in End-to-End Speaker Verification System2018
Barlow Twins self-supervised learning for robust speaker recognition2022
Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems2020
Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement2020
RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker Verification2019