VoxCeleb: A Large-Scale Speaker Identification Dataset
VoxCeleb: крупномасштабный набор данных для идентификации говорящих
2017-08-16
SCID: 54.1/7urkj3bv
Discuss with AI
VoxCeleb datasetactive speaker verificationfacial recognitionspeaker identificationtwo-stream synchronization CNN
Figures from the paper
Abstract (AI)
Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent speaker identification dataset collected 'in the wild'. We make two contributions. First, we propose a fully automated pipeline based on computer vision techniques to create the dataset from open-source media. Our pipeline involves obtaining videos from YouTube; performing active speaker verification using a two-stream synchronization Convolutional Neural Network (CNN), and confirming the identity of the speaker using CNN based facial recognition. We use this pipeline to curate VoxCeleb which contains hundreds of thousands of 'real world' utterances for over 1,000 celebrities. Our second contribution is to apply and compare various state of the art speaker identification techniques on our dataset to establish baseline performance. We show that a CNN based architecture obtains the best performance for both identification and verification.
Key Findings
1
A fully automated pipeline combines video acquisition, two-stream CNN active-speaker verification, and CNN-based facial recognition to curate and confirm speaker samples.
2
CNN-based architectures achieve the best performance among evaluated methods for both speaker identification and speaker verification.
3
The dataset is collected from YouTube videos under unconstrained, in-the-wild conditions, addressing the limited scale and constrained environments of existing datasets.
4
The paper compares state-of-the-art speaker identification methods on VoxCeleb and establishes baseline performance for identification and verification.
5
VoxCeleb is a large-scale, text-independent speaker identification dataset containing hundreds of thousands of real-world utterances from over 1,000 celebrities.
Research Object
VoxCeleb large-scale text-independent speaker identification dataset, comprising real-world speech utterances from over 1,000 celebrities collected from open-source media
Research Subject
Speaker identification and verification performance on in-the-wild speech, including the comparative effectiveness of state-of-the-art methods and CNN-based architectures
Publication Details
Publication Date
2017-08-16
Journal
Publisher
ISSN
Cited by
2178
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by10
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing2022
Voxceleb: Large-scale speaker verification in the wild2019
Self-Supervised Speech Representation Learning: A Review2022
Speaker Normalization for Self-Supervised Speech Emotion Recognition2022
Self-Supervised Speaker Recognition with Loss-Gated Learning2022
End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-Entropy2022
Analysis of Length Normalization in End-to-End Speaker Verification System2018
Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of any Number of Speakers2020
RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker Verification2019