Voxceleb: Large-scale speaker verification in the wild
VoxCeleb: Масштабная верификация говорящих в естественных условиях
2019-10-16
SCID: 54.1/a5qj9jpg
Discuss with AI
VoxCeleb datasetactive speaker verificationspeaker recognition in the wildspeaker verificationtwo-stream synchronization CNN
Figures from the paper
Abstract (AI)
The objective of this work is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual dataset collected from open source media using a fully automated pipeline. Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and usually require manual annotations, hence are limited in size. We propose a pipeline based on computer vision techniques to create the dataset from open-source media. Our pipeline involves obtaining videos from YouTube; performing active speaker verification using a two-stream synchronization Convolutional Neural Network (CNN), and confirming the identity of the speaker using CNN based facial recognition. We use this pipeline to curate VoxCeleb which contains contains over a million ‘real-world’ utterances from over 6000 speakers. This is several times larger than any publicly available speaker recognition dataset. Second, we develop and compare different CNN architectures with various aggregation methods and training loss functions that can effectively recognise identities from voice under various conditions. The models trained on our dataset surpass the performance of previous works by a significant margin.
Key Findings
1
Evaluates CNN architectures, aggregation methods, and training losses for speaker identification under varied conditions.
2
Introduces VoxCeleb, an automatically curated audio-visual dataset containing over one million real-world utterances from more than 6,000 speakers.
3
Models trained on VoxCeleb significantly outperform previous approaches in speaker-recognition performance.
4
Uses active-speaker verification with a two-stream synchronization CNN and CNN-based facial recognition to construct the dataset from YouTube videos.
5
VoxCeleb is several times larger than previously available public speaker-recognition datasets and captures noisy, unconstrained conditions.
Research Object
speaker recognition in noisy and unconstrained real-world conditions
Research Subject
voice-based speaker identity recognition performance across varying conditions
Publication Details
Publication Date
2019-10-16
Journal
Publisher
ISSN
Cited by
686
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai8
ImageNet classification with deep convolutional neural networks2017
Front-End Factor Analysis for Speaker Verification2010
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
VoxCeleb: A Large-Scale Speaker Identification Dataset2017
Generalized End-to-End Loss for Speaker Verification2018
End-to-end text-dependent speaker verification2016
Analysis of Length Normalization in End-to-End Speaker Verification System2018
Deep Residual Learning for Image Recognition2016
Cited by4
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing2022
SSAST: Self-Supervised Audio Spectrogram Transformer2022
Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?2022
Barlow Twins self-supervised learning for robust speaker recognition2022