End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-Entropy
Сквозная верификация дикторов с использованием взвешенной бинарной кросс-энтропии на основе бипартитного ранжирования и обучения по принципу усложнения задач
2022-01-01
SCID: 54.1/f3krsnct
Discuss with AI
bipartite rankingcurriculum learningend-to-end speaker verificationtrial imbalanceweighted binary cross-entropy
Figures from the paper
Abstract (AI)
End-to-end speaker verification achieves the verification through estimating directly the similarity score between a pair of utterances, which is formulated as a binary (i.e., target versus non-target) classification problem. Unlike the stage-wise method, an end-to-end verification approach optimizes the evaluation metrics directly and its output layer is parameter-free, which can save great computing and memory resources. However, there are two important issues that need to be meticulously handled in training an end-to-end speaker verification model. The first one is how to deal with severely imbalanced trials, i.e., the number of target trials is much smaller than that of nontarget trials, and the other is about how to handle easy trials that do not help improve the model in training. To circumvent these two issues, we propose in this paper a binary cross-entropy (BCE) type of loss function and present a method to train the deep neural network (DNN) models based on the proposed loss function for end-to-end speaker verification. The training process employs a bipartite ranking method to deal with the trial imbalance problem and a curriculum learning method to help improve both the training stability and performance of the model by selecting non-target trials from easy to hard ones gradually along the convergence process. Since the training process employs bipartite ranking and curriculum learning and the loss function is of the generalized BCE form, we name the new approach \textit{curriculum bipartite ranking weighted binary cross-entropy} (CBRW-BCE). Experimental results show that the model trained with CBRW-BCE not only achieves the state-of-the-art performance but is also well calibrated.
Key Findings
1
Bipartite ranking addresses severe imbalance between target and non-target verification trials during training.
2
Curriculum learning gradually selects non-target trials from easy to hard, improving training stability and model performance.
3
The end-to-end formulation directly optimizes verification metrics and uses a parameter-free output layer, reducing computing and memory requirements compared with stage-wise methods.
4
The paper introduces CBRW-BCE, a generalized binary cross-entropy loss for end-to-end speaker verification.
5
The proposed model achieves state-of-the-art speaker-verification performance and produces well-calibrated scores.
Research Object
end-to-end speaker verification models operating on pairs of target and non-target utterances
Research Subject
the effects of severely imbalanced and easy trials on training, and the resulting verification performance, stability, and calibration
Publication Details
Publication Date
2022-01-01
Journal
Publisher
ISSN
Cited by
37
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai6
FaceNet: A unified embedding for face recognition and clustering2015
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
VoxCeleb: A Large-Scale Speaker Identification Dataset2017
A study on data augmentation of reverberant speech for robust speech recognition2017
Generalized End-to-End Loss for Speaker Verification2018
End-to-end text-dependent speaker verification2016