Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of any Number of Speakers

Совместный подсчёт говорящих, распознавание речи и идентификация говорящих для перекрывающейся речи с произвольным числом говорящих
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Tianyan Zhou, Takuya Yoshioka
2020-10-25

overlapped speech recognitionserialized output trainingspeaker identificationspeaker-attributed automatic speech recognitionspeaker-attributed word error rate
We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech.Our model is built on serialized output training (SOT) with attention-based encoder-decoder, a recently proposed method for recognizing overlapped speech comprising an arbitrary number of speakers.We extend SOT by introducing a speaker inventory as an auxiliary input to produce speaker labels as well as multi-speaker transcriptions.All model parameters are optimized by speaker-attributed maximum mutual information criterion, which represents a joint probability for overlapped speech recognition and speaker identification.Experiments on LibriSpeech corpus show that our proposed method achieves significantly better speaker-attributed word error rate than the baseline that separately performs overlapped speech recognition and speaker identification.
1
All model parameters are optimized with a speaker-attributed maximum mutual information criterion jointly modeling overlapped-speech recognition and speaker identification.
2
An end-to-end model jointly performs speaker counting, speech recognition, and speaker identification for monaural overlapped speech with any number of speakers.
3
On the LibriSpeech corpus, the proposed approach achieves significantly lower speaker-attributed word error rate than a baseline that separates recognition and speaker identification.
4
The method extends serialized output training by using a speaker inventory as an auxiliary input to generate speaker labels alongside multi-speaker transcriptions.

monaural overlapped speech with an arbitrary number of speakers

joint speaker counting, multi-speaker speech recognition, and speaker identification, evaluated by speaker-attributed word error rate

Publication Details
Publication Date
2020-10-25
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Naoyuki Kanda
Yashesh Gaur
Xiaofei Wang
Zhong Meng
Zhuo Chen
Tianyan Zhou
Takuya Yoshioka
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%