Deep neural network-based speaker embeddings for end-to-end speaker verification

Векторные представления говорящих на основе глубокой нейронной сети для сквозной верификации говорящего
David Snyder, Pegah Ghahremani, Daniel Povey, Daniel Garcia-Romero, Yishay Carmiel, Sanjeev Khudanpur
2016-12-01

deep neural networksend-to-end speaker verificationequal error ratespeaker embeddingstext-independent speaker verification
In this study, we investigate an end-to-end text-independent speaker verification system. The architecture consists of a deep neural network that takes a variable length speech segment and maps it to a speaker embedding. The objective function separates same-speaker and different-speaker pairs, and is reused during verification. Similar systems have recently shown promise for text-dependent verification, but we believe that this is unexplored for the text-independent task. We show that given a large number of training speakers, the proposed system outperforms an i-vector baseline in equal error-rate (EER) and at low miss rates. Relative to the baseline, the end-to-end system reduces EER by 13% average and 29% pooled across test conditions. The fused system achieves a reduction of 32% average and 38% pooled.
1
An end-to-end text-independent speaker verification system maps variable-length speech segments directly to speaker embeddings using a deep neural network.
2
Compared with the baseline, the end-to-end system reduces EER by 13% on average and 29% when pooled across test conditions.
3
Fusing systems further reduces EER by 32% on average and 38% when pooled across test conditions.
4
The training objective explicitly separates same-speaker and different-speaker pairs and is reused during verification.
5
With many training speakers, the proposed system outperforms an i-vector baseline in equal error rate and at low miss rates.

end-to-end text-independent speaker verification system

verification performance and speaker-discriminative embedding behavior, including equal error rate and low-miss-rate performance relative to an i-vector baseline

Publication Details
Publication Date
2016-12-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
David Snyder
Pegah Ghahremani
Daniel Povey
Daniel Garcia-Romero
Yishay Carmiel
Sanjeev Khudanpur
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%