Self-Supervised Speech Representation Learning: A Review

Обучение речевых представлений с самоконтролем: обзор
Hung-yi Lee, Tara N. Sainath, Christian Igel, Karen Livescu, Shinji Watanabe, Shang-Wen Li, Jakob D. Havtorn, Lasse Borgholt, Lars Maaløe, Katrin Kirchhoff, Abdelrahman Mohamed, Joakim Edin
2022-09-15

automatic speech recognitioncontrastive learninggenerative methodspredictive methodsself-supervised speech representation learning
Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply this to dialects and languages for which only limited labeled data is available. Self-supervised representation learning methods promise a single universal model that would benefit a wide variety of tasks and domains. Such methods have shown success in natural language processing and computer vision domains, achieving new levels of performance while reducing the number of labels required for many downstream scenarios. Speech representation learning is experiencing similar progress in three main categories: generative, contrastive, and predictive methods. Other approaches rely on multi-modal data for pre-training, mixing text or visual data streams with speech. Although self-supervised speech representation is still a nascent research area, it is closely related to acoustic word embedding and learning with zero lexical resources, both of which have seen active research for many years. This review presents approaches for self-supervised speech representation learning and their connection to other research areas. Since many current methods focus solely on automatic speech recognition as a downstream task, we review recent efforts on benchmarking learned representations to extend the application beyond speech recognition.
1
Because many methods are evaluated mainly on automatic speech recognition, recent benchmarking efforts assess their usefulness for broader applications.
2
Current approaches primarily fall into generative, contrastive, and predictive categories, with some incorporating multimodal speech, text, or visual pre-training.
3
Self-supervised speech representation learning aims to produce universal models usable across diverse speech tasks and domains.
4
Self-supervised speech representation learning is connected to acoustic word embedding and zero-lexical-resource research.
5
These methods can reduce dependence on labeled data, addressing challenges for low-resource languages and dialects.

self-supervised speech representations and their learning methods

the approaches, categories, cross-modal connections, and downstream-task generalization of self-supervised speech representation learning

Publication Details
Publication Date
2022-09-15
Journal
Publisher
ISSN
Cited by
370
Access Type
Author Information
Authors
Hung-yi Lee
Tara N. Sainath
Christian Igel
Karen Livescu
Shinji Watanabe
Shang-Wen Li
Jakob D. Havtorn
Lasse Borgholt
Lars Maaløe
Katrin Kirchhoff
Abdelrahman Mohamed
Joakim Edin
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%