Front-End Factor Analysis for Speaker Verification
Факторный анализ на переднем конце для верификации диктора
2010-08-10
SCID: 54.1/6354mjmg
Discuss with AI
channel compensationcosine scoringfactor analysisspeaker verificationtotal variability space
Figures from the paper
Abstract (AI)
This paper presents an extension of our previous work which proposes a new speaker representation for speaker verification. In this modeling, a new low-dimensional speaker- and channel-dependent space is defined using a simple factor analysis. This space is named the total variability space because it models both speaker and channel variabilities. Two speaker verification systems are proposed which use this new representation. The first system is a support vector machine-based system that uses the cosine kernel to estimate the similarity between the input data. The second system directly uses the cosine similarity as the final decision score. We tested three channel compensation techniques in the total variability space, which are within-class covariance normalization (WCCN), linear discriminate analysis (LDA), and nuisance attribute projection (NAP). We found that the best results are obtained when LDA is followed by WCCN. We achieved an equal error rate (EER) of 1.12% and MinDCF of 0.0094 using the cosine distance scoring on the male English trials of the core condition of the NIST 2008 Speaker Recognition Evaluation dataset. We also obtained 4% absolute EER improvement for both-gender trials on the 10 s-10 s condition compared to the classical joint factor analysis scoring.
Key Findings
1
Among tested channel-compensation methods, applying linear discriminant analysis followed by within-class covariance normalization produced the best results.
2
Cosine-distance scoring achieved 1.12% EER and 0.0094 MinDCF on male English trials in the NIST 2008 core condition.
3
The paper introduces a low-dimensional total variability space modeling both speaker and channel variability through simple factor analysis.
4
The proposed approach improved absolute EER by 4% on both-gender 10-second enrollment and test trials compared with classical joint factor-analysis scoring.
5
Two verification systems use the new representation: cosine-kernel support vector machines and direct cosine-similarity scoring.
Research Object
speaker verification systems using low-dimensional total variability speaker representations
Research Subject
the effects of speaker and channel variability modeling and channel compensation on verification performance, including cosine-s similarity scoring, EER, and MinDCF
Publication Details
Publication Date
2010-08-10
Journal
Publisher
ISSN
Cited by
3602
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by10
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
A Comparative Evaluation of Unsupervised Anomaly Detection Algorithms for Multivariate Data2016
Generalized End-to-End Loss for Speaker Verification2018
Voxceleb: Large-scale speaker verification in the wild2019
End-to-end text-dependent speaker verification2016
Self-Supervised Speech Representation Learning: A Review2022
RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker Verification2019
Deep neural network-based speaker embeddings for end-to-end speaker verification2016
End-to-End attention based text-dependent speaker verification2016
Source-Normalized LDA for Robust Speaker Recognition Using i-Vectors From Multiple Speech Sources2011