WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
WavLM: крупномасштабное самосупервизорное предварительное обучение для комплексной обработки речи
2022-07-04
SCID: 54.1/sy3puq9b
Discuss with AI
SUPERB benchmarkWavLMmasked speech predictionself-supervised speech pre-trainingspeech denoising
Figures from the paper
Abstract (AI)
Self-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks. As speech signal contains multi-faceted information including speaker identity, paralinguistics, spoken content, etc., learning universal representations for all speech tasks is challenging. To tackle the problem, we propose a new pre-trained model, WavLM, to solve full-stack downstream speech tasks. WavLM jointly learns masked speech prediction and denoising in pre-training. By this means, WavLM does not only keep the speech content modeling capability by the masked speech prediction, but also improves the potential to non-ASR tasks by the speech denoising. In addition, WavLM employs gated relative position bias for the Transformer structure to better capture the sequence ordering of input speech. We also scale up the training dataset from 60 k hours to 94 k hours. WavLM Large achieves state-of-the-art performance on the SUPERB benchmark, and brings significant improvements for various speech processing tasks on their representative benchmarks.
Key Findings
1
Joint masked speech prediction and denoising preserves speech-content modeling while improving representations for non-ASR tasks involving speaker and paralinguistic information.
2
The pre-training dataset is scaled from 60,000 to 94,000 hours of speech, supporting larger-scale representation learning.
3
WavLM Large achieves state-of-the-art performance on the SUPERB benchmark and significant improvements across representative speech-processing benchmarks.
4
WavLM introduces gated relative position bias in its Transformer architecture to better capture the temporal ordering of speech inputs.
5
WavLM is a self-supervised pre-trained model designed to learn universal representations for full-stack speech processing tasks beyond automatic speech recognition.
Research Object
WavLM pre-trained speech representations for full-stack downstream speech processing tasks
Research Subject
The capability and performance of WavLM representations in modeling speech content and non-ASR information across diverse speech processing tasks
Publication Details
Publication Date
2022-07-04
Journal
Publisher
ISSN
Cited by
1863
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai9
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
Representation Learning with Contrastive Predictive Coding2018
VoxCeleb: A Large-Scale Speaker Identification Dataset2017
A study on data augmentation of reverberant speech for robust speech recognition2017
Voxceleb: Large-scale speaker verification in the wild2019
TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech2021
SSAST: Self-Supervised Audio Spectrogram Transformer2022