Self-Supervised Speech Representation Learning: A Review
Обучение речевых представлений с самоконтролем: обзор
2022-09-15
SCID: 54.1/zvuczv9w
Discuss with AI
automatic speech recognitioncontrastive learninggenerative methodspredictive methodsself-supervised speech representation learning
Figures from the paper
Abstract (AI)
Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply this to dialects and languages for which only limited labeled data is available. Self-supervised representation learning methods promise a single universal model that would benefit a wide variety of tasks and domains. Such methods have shown success in natural language processing and computer vision domains, achieving new levels of performance while reducing the number of labels required for many downstream scenarios. Speech representation learning is experiencing similar progress in three main categories: generative, contrastive, and predictive methods. Other approaches rely on multi-modal data for pre-training, mixing text or visual data streams with speech. Although self-supervised speech representation is still a nascent research area, it is closely related to acoustic word embedding and learning with zero lexical resources, both of which have seen active research for many years. This review presents approaches for self-supervised speech representation learning and their connection to other research areas. Since many current methods focus solely on automatic speech recognition as a downstream task, we review recent efforts on benchmarking learned representations to extend the application beyond speech recognition.
Key Findings
1
Because many methods are evaluated mainly on automatic speech recognition, recent benchmarking efforts assess their usefulness for broader applications.
2
Current approaches primarily fall into generative, contrastive, and predictive categories, with some incorporating multimodal speech, text, or visual pre-training.
3
Self-supervised speech representation learning aims to produce universal models usable across diverse speech tasks and domains.
4
Self-supervised speech representation learning is connected to acoustic word embedding and zero-lexical-resource research.
5
These methods can reduce dependence on labeled data, addressing challenges for low-resource languages and dialects.
Research Object
self-supervised speech representations and their learning methods
Research Subject
the approaches, categories, cross-modal connections, and downstream-task generalization of self-supervised speech representation learning
Publication Details
Publication Date
2022-09-15
Journal
Publisher
ISSN
Cited by
370
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai14
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Reducing the Dimensionality of Data with Neural Networks2006
Representation Learning: A Review and New Perspectives2013
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups2012
Machine learning: Trends, perspectives, and prospects2015
Emerging Properties in Self-Supervised Vision Transformers2021
Representation Learning with Contrastive Predictive Coding2018
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing2022
Front-End Factor Analysis for Speaker Verification2010
On the Opportunities and Risks of Foundation Models2021
VoxCeleb: A Large-Scale Speaker Identification Dataset2017
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing2022
Efficient Transformers: A Survey2022
TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech2021