Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker Embeddings

Адаптация моделей самоконтролируемого обучения для распознавания речи нескольких говорящих с использованием эмбеддингов говорящих
Zili Huang, Desh Raj, Leibny Paola Garcia, Sanjeev Khudanpur
2023-05-05

joint speaker modelingmulti-talker speech recognitionself-supervised learningspeaker embeddingstarget speaker extraction
Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have degraded performance for multi-talker scenarios — possibly due to the domain mismatch — which severely limits their use for such applications. In this paper, we investigate the adaptation of upstream SSL models to the multi-talker automatic speech recognition (ASR) task under two conditions. First, when segmented utterances are given, we show that adding a target speaker extraction (TSE) module based on enrollment embeddings is complementary to mixture-aware pre-training. Second, for unsegmented mixtures, we propose a novel joint speaker modeling (JSM) approach, which aggregates information from all speakers in the mixture through their embeddings. With controlled experiments on Libri2Mix, we show that using speaker embeddings provides relative WER improvements of 9.1% and 42.1% over strong baselines for the segmented and unsegmented cases, respectively. We also demonstrate the effectiveness of our models for real conversational mixtures through experiments on the AMI dataset. Our code and models are open-sourced on https://github.com/HuangZiliAndy/SSL_for_multitalker.
1
For segmented utterances, adding a target speaker extraction module based on enrollment embeddings complements mixture-aware pre-training.
2
For unsegmented mixtures, the proposed joint speaker modeling approach aggregates information from all mixture speakers through their embeddings.
3
On Libri2Mix, speaker embeddings improve relative WER by 9.1% for segmented mixtures and 42.1% for unsegmented mixtures over strong baselines.
4
Self-supervised learning models show degraded automatic speech recognition performance in multi-talker scenarios, likely because of domain mismatch.
5
The proposed models also demonstrate effectiveness on real conversational mixtures from the AMI dataset, with code and models publicly released.

upstream self-supervised speech models adapted to multi-talker automatic speech recognition of segmented utterances and unsegmented speech mixtures

the impact of speaker-embedding-based target speaker extraction and joint speaker modeling on multi-talker ASR performance

Publication Details
Publication Date
2023-05-05
Journal
Publisher
ISSN
Cited by
29
Access Type
Author Information
Authors
Zili Huang
Desh Raj
Leibny Paola Garcia
Sanjeev Khudanpur
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%