WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

WavLM: крупномасштабное самосупервизорное предварительное обучение для комплексной обработки речи
Furu Wei, Michael Zeng, Shujie Liu, Jinyu Li, Xiangzhan Yu, Yao Qian, Zhuo Chen, Yu Wu, Long Zhou, Chengyi Wang, Shuo Ren, Takuya Yoshioka, Sanyuan Chen, Zhengyang Chen, Naoyuki Kanda, Xiong Xiao, Jian Wu, Yanmin Qian
2022-07-04

SUPERB benchmarkWavLMmasked speech predictionself-supervised speech pre-trainingspeech denoising
Self-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks. As speech signal contains multi-faceted information including speaker identity, paralinguistics, spoken content, etc., learning universal representations for all speech tasks is challenging. To tackle the problem, we propose a new pre-trained model, WavLM, to solve full-stack downstream speech tasks. WavLM jointly learns masked speech prediction and denoising in pre-training. By this means, WavLM does not only keep the speech content modeling capability by the masked speech prediction, but also improves the potential to non-ASR tasks by the speech denoising. In addition, WavLM employs gated relative position bias for the Transformer structure to better capture the sequence ordering of input speech. We also scale up the training dataset from 60 k hours to 94 k hours. WavLM Large achieves state-of-the-art performance on the SUPERB benchmark, and brings significant improvements for various speech processing tasks on their representative benchmarks.
1
Joint masked speech prediction and denoising preserves speech-content modeling while improving representations for non-ASR tasks involving speaker and paralinguistic information.
2
The pre-training dataset is scaled from 60,000 to 94,000 hours of speech, supporting larger-scale representation learning.
3
WavLM Large achieves state-of-the-art performance on the SUPERB benchmark and significant improvements across representative speech-processing benchmarks.
4
WavLM introduces gated relative position bias in its Transformer architecture to better capture the temporal ordering of speech inputs.
5
WavLM is a self-supervised pre-trained model designed to learn universal representations for full-stack speech processing tasks beyond automatic speech recognition.

WavLM pre-trained speech representations for full-stack downstream speech processing tasks

The capability and performance of WavLM representations in modeling speech content and non-ASR information across diverse speech processing tasks

Publication Details
Publication Date
2022-07-04
Journal
Publisher
ISSN
Cited by
1863
Access Type
Author Information
Authors
Furu Wei
Michael Zeng
Shujie Liu
Jinyu Li
Xiangzhan Yu
Yao Qian
Zhuo Chen
Yu Wu
Long Zhou
Chengyi Wang
Shuo Ren
Takuya Yoshioka
Sanyuan Chen
Zhengyang Chen
Naoyuki Kanda
Xiong Xiao
Jian Wu
Yanmin Qian
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%