End-to-end text-dependent speaker verification
Сквозная текстозависимая верификация говорящего
2016-03-01
SCID: 54.1/nxz4kv6e
Discuss with AI
Ok Google benchmarkend-to-end neural networkfew-shot speaker verificationspeaker model estimationtext-dependent speaker verification
Figures from the paper
Abstract (AI)
In this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the system's components using the same evaluation protocol and metric as at test time. Such an approach will result in simple and efficient systems, requiring little domain-specific knowledge and making few model assumptions. We implement the idea by formulating the problem as a single neural network architecture, including the estimation of a speaker model on only a few utterances, and evaluate it on our internal "Ok Google" benchmark for text-dependent speaker verification. The proposed approach appears to be very effective for big data applications Like ours that require highly accurate, easy-to-maintain systems with a small footprint.
Key Findings
1
A single neural-network architecture estimates a speaker model from only a few reference utterances, reducing reliance on domain-specific knowledge and model assumptions.
2
All system components are jointly optimized using the same evaluation protocol and metric applied during testing.
3
Evaluation on an internal “Ok Google” text-dependent speaker-verification benchmark indicates the approach is highly effective for accurate, maintainable, small-footprint systems.
4
The approach is designed for large-scale applications and emphasizes simple, efficient deployment rather than complex separately engineered components.
5
The paper introduces an end-to-end, data-driven speaker-verification system that maps a test utterance and a few reference utterances directly to one verification score.
Research Object
text-dependent speaker verification system
Research Subject
end-to-end data-driven verification performance, including accurate score estimation from a test utterance and a few reference utterances
Publication Details
Publication Date
2016-03-01
Journal
Publisher
ISSN
Cited by
533
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai2
Cited by8
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
Generalized End-to-End Loss for Speaker Verification2018
Voxceleb: Large-scale speaker verification in the wild2019
End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-Entropy2022
Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement2020
A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result2018
Deep neural network-based speaker embeddings for end-to-end speaker verification2016
End-to-End attention based text-dependent speaker verification2016