CAMSAT: Augmentation Mix and Self-Augmented Training Clustering for Self-Supervised Speaker Recognition

CAMSAT: кластеризация с объединением аугментаций и самоаугментированным обучением для самообучаемого распознавания говорящих
Abderrahim Fathan, Jahangir Alam
2023-12-16

Self-Augmented Trainingpseudo-label clusteringself-supervised speaker recognitionspeaker embeddingsspeaker verification
Clustering (CL)-based pseudo-labels (PLs) are widely used to optimize speaker embedding (SE) networks and train self-supervised (SS) speaker verification (SV) systems. However, PL-based SS training depends on high-quality PLs. In this paper, we propose a general-purpose CL algorithm called CAMSAT that outperforms all other baselines used to cluster SEs. Moreover, using the generated PLs to train our SE system allows us to further improve SV performance. CAMSAT is based on two principles: (1) mixing predictions of augmented samples to provide a complementary supervisory signal for CL and enforce symmetry within augmentations (2) Self-Augmented Training to enforce representation invariance and maximize the information-theoretic dependency between samples and their predicted PLs. We provide a thorough comparative analysis of the performance of our CL method vs. all baselines using a variety of CL metrics and perform an ablation study to analyze the contribution of each component.
1
CAMSAT is a general-purpose clustering algorithm for speaker embeddings that outperforms all evaluated clustering baselines.
2
CAMSAT mixes predictions from augmented samples, providing complementary supervision and enforcing consistency across augmentations during clustering.
3
Its Self-Augmented Training component promotes representation invariance and increases the information-theoretic dependency between samples and predicted pseudo-labels.
4
Pseudo-labels generated by CAMSAT improve speaker verification performance when used to train the speaker embedding system.
5
The study compares CAMSAT with baselines across multiple clustering metrics and uses ablation experiments to assess each component’s contribution.

self-supervised speaker recognition systems and speaker embeddings

clustering-based pseudo-label generation and its effect on speaker-verification performance, including augmentation consistency and representation invariance

Publication Details
Publication Date
2023-12-16
Journal
Publisher
ISSN
Cited by
5
Access Type
Author Information
Authors
Abderrahim Fathan
Jahangir Alam
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%