End-to-end text-dependent speaker verification

Сквозная текстозависимая верификация говорящего
Noam Shazeer, Georg Heigold, Samy Bengio, Ignacio López Moreno
2016-03-01

Ok Google benchmarkend-to-end neural networkfew-shot speaker verificationspeaker model estimationtext-dependent speaker verification
In this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the system's components using the same evaluation protocol and metric as at test time. Such an approach will result in simple and efficient systems, requiring little domain-specific knowledge and making few model assumptions. We implement the idea by formulating the problem as a single neural network architecture, including the estimation of a speaker model on only a few utterances, and evaluate it on our internal "Ok Google" benchmark for text-dependent speaker verification. The proposed approach appears to be very effective for big data applications Like ours that require highly accurate, easy-to-maintain systems with a small footprint.
1
A single neural-network architecture estimates a speaker model from only a few reference utterances, reducing reliance on domain-specific knowledge and model assumptions.
2
All system components are jointly optimized using the same evaluation protocol and metric applied during testing.
3
Evaluation on an internal “Ok Google” text-dependent speaker-verification benchmark indicates the approach is highly effective for accurate, maintainable, small-footprint systems.
4
The approach is designed for large-scale applications and emphasizes simple, efficient deployment rather than complex separately engineered components.
5
The paper introduces an end-to-end, data-driven speaker-verification system that maps a test utterance and a few reference utterances directly to one verification score.

text-dependent speaker verification system

end-to-end data-driven verification performance, including accurate score estimation from a test utterance and a few reference utterances

Publication Details
Publication Date
2016-03-01
Journal
Publisher
ISSN
Cited by
533
Access Type
Author Information
Authors
Noam Shazeer
Georg Heigold
Samy Bengio
Ignacio López Moreno
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%