A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result
Полная сквозная система верификации говорящего на основе глубоких нейронных сетей: от исходных сигналов до результата верификации
2018-04-01
SCID: 54.1/7mgez8yy
Discuss with AI
deep neural networksend-to-end speaker verificationlong short-term memoryraw audio signalsstrided convolution
Figures from the paper
Abstract (AI)
End-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signals have not been explored in speaker verification. In this paper, a complete end-to-end speaker verification system is proposed, which inputs raw audio signals and outputs the verification results. A pre-processing layer and the embedded speaker feature extraction models were mainly investigated. The proposed pre-emphasis layer was combined with a strided convolution layer for pre-processing at the first two hidden layers. In addition, speaker feature extraction models using convolutionallayer and long short-term memory are proposed to be embedded in the proposed end-to-end system.
Key Findings
1
A learned pre-emphasis layer combined with strided convolution is used for raw-signal preprocessing in the first two hidden layers.
2
Convolutional and long short-term memory models are developed as embedded speaker feature extractors within the end-to-end architecture.
3
The paper proposes a complete end-to-end speaker verification system that directly maps raw audio signals to verification results.
4
The work addresses the previously unexplored use of raw audio inputs in end-to-end speaker verification, according to the abstract.
Research Object
End-to-end speaker verification system processing raw audio signals
Research Subject
Pre-processing and embedded speaker-feature extraction architectures, including pre-emphasis with strided convolution and convolutional/LSTM models, for producing verification results
Publication Details
Publication Date
2018-04-01
Journal
Publisher
ISSN
Access Type
Author Information
Download PDF
Subscribe to digest