RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker Verification

RawNet: усовершенствованная сквозная глубокая нейронная сеть, использующая исходные речевые сигналы для текстонезависимой верификации говорящего
Jee-weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-jin Shim, Ha-Jin Yu
2019-09-13

End-to-end deep neural networksRaw waveform modelingSpeaker embeddingsSpeaker verificationVoxCeleb1
Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains.In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring further investigation.In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional objective functions, and back-end classification.Adjustment of model architecture using a pre-training scheme can extract speaker embeddings, giving a significant improvement in performance.Additional objective functions simplify the process of extracting speaker embeddings by merging conventional two-phase processes: extracting utterance-level features such as i-vectors or x-vectors and the feature enhancement phase, e.g., linear discriminant analysis.Effective back-end classification models that suit the proposed speaker embedding are also explored.We propose an end-toend system that comprises two deep neural networks, one frontend for utterance-level speaker embedding extraction and the other for back-end classification.Experiments conducted on the VoxCeleb1 dataset demonstrate that the proposed model achieves state-of-the-art performance among systems without data augmentation.The proposed system is also comparable to the state-of-the-art x-vector system that adopts data augmentation.
1
Additional objective functions integrate speaker embedding extraction and feature enhancement, replacing conventional two-stage i-vector or x-vector pipelines with a unified process.
2
Architecture adjustment combined with pre-training substantially improves raw-waveform speaker embedding extraction.
3
On VoxCeleb1, RawNet achieves state-of-the-art performance among systems without data augmentation and performs comparably to an augmented x-vector system.
4
RawNet directly models raw waveforms with an end-to-end deep neural network for text-independent speaker verification.
5
The proposed system uses separate front-end and back-end neural networks for utterance-level speaker embedding extraction and speaker classification.

end-to-end deep neural network systems using raw waveforms for text-independent speaker verification

speaker embedding extraction and back-end classification performance, including the effects of model architecture, pre-training, additional objective functions, and classification models

Publication Details
Publication Date
2019-09-13
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Jee-weon Jung
Hee-Soo Heo
Ju-ho Kim
Hye-jin Shim
Ha-Jin Yu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%