Self-Supervised Speaker Recognition with Loss-Gated Learning
Самостоятельно контролируемое распознавание говорящих с обучением, управляемым функцией потерь
2022-04-27
SCID: 54.1/jbybn4b2
Discuss with AI
VoxCeleb2equal error rateloss-gated learningpseudo labelsself-supervised speaker recognition
Figures from the paper
Abstract (AI)
In self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to model the data with reliable labels faster than those with unreliable labels. This motivates us to study a loss-gated learning (LGL) strategy, which extracts the reliable labels through the fitting ability of the neural network during training. With the proposed LGL, our speaker recognition model obtains a 46.3% performance gain over the system without it. Further, the proposed self-supervised speaker recognition with LGL trained on the VoxCeleb2 dataset without any labels achieves an equal error rate of 1.66% on the VoxCeleb1 original test set.
Key Findings
1
A fully self-supervised model trained on unlabeled VoxCeleb2 achieves a 1.66% equal error rate on the VoxCeleb1 original test set.
2
Loss-gated learning improves speaker recognition performance by 46.3% compared with a system without loss gating.
3
Speaker recognition networks fit samples with reliable pseudo labels faster than samples with unreliable pseudo labels.
4
The proposed loss-gated learning strategy identifies reliable pseudo labels based on the network’s fitting behavior during training.
Research Object
self-supervised speaker recognition model trained with pseudo labels
Research Subject
the effect of pseudo-label reliability and loss-gated learning on speaker-recognition performance
Publication Details
Publication Date
2022-04-27
Journal
Publisher
ISSN
Cited by
59
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai6
Adam: A Method for Stochastic Optimization2014
Adam: A Method for Stochastic Optimization2015
Least squares quantization in PCM1982
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
VoxCeleb: A Large-Scale Speaker Identification Dataset2017
A study on data augmentation of reverberant speech for robust speech recognition2017