Racial disparities in automated speech recognition

Расовые различия в автоматическом распознавании речи
Dan Jurafsky, Joe Nudell, Sharad Goel, Allison Koenecke, Andrew Nam, Emily Lake, Minnie Quartey, Zion Mengesha, Connor Toups, John R. Rickford
2020-03-23

African American Vernacular Englishacoustic modelsautomated speech recognitionracial disparitiesword error rate
Automated speech recognition (ASR) systems, which use sophisticated machine-learning algorithms to convert spoken language to text, have become increasingly widespread, powering popular virtual assistants, facilitating automated closed captioning, and enabling digital dictation platforms for health care. Over the last several years, the quality of these systems has dramatically improved, due both to advances in deep learning and to the collection of large-scale datasets used to train the systems. There is concern, however, that these tools do not work equally well for all subgroups of the population. Here, we examine the ability of five state-of-the-art ASR systems-developed by Amazon, Apple, Google, IBM, and Microsoft-to transcribe structured interviews conducted with 42 white speakers and 73 black speakers. In total, this corpus spans five US cities and consists of 19.8 h of audio matched on the age and gender of the speaker. We found that all five ASR systems exhibited substantial racial disparities, with an average word error rate (WER) of 0.35 for black speakers compared with 0.19 for white speakers. We trace these disparities to the underlying acoustic models used by the ASR systems as the race gap was equally large on a subset of identical phrases spoken by black and white individuals in our corpus. We conclude by proposing strategies-such as using more diverse training datasets that include African American Vernacular English-to reduce these performance differences and ensure speech recognition technology is inclusive.
1
All five systems exhibited substantial racial disparities, with an average word error rate of 0.35 for Black speakers versus 0.19 for white speakers.
2
Five state-of-the-art ASR systems from Amazon, Apple, Google, IBM, and Microsoft were evaluated using 19.8 hours of interviews from 115 speakers across five US cities.
3
The authors recommend more diverse training datasets, including African American Vernacular English, to reduce racial disparities in speech recognition.
4
The performance gap persisted for identical phrases spoken by Black and white participants, implicating differences in the systems’ underlying acoustic models.

Five state-of-the-art automated speech recognition (ASR) systems transcribing structured interviews by Black and white U.S. speakers

Racial disparities in ASR transcription performance, measured by word error rates and linked to the systems’ acoustic models

Publication Details
Publication Date
2020-03-23
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Dan Jurafsky
Joe Nudell
Sharad Goel
Allison Koenecke
Andrew Nam
Emily Lake
Minnie Quartey
Zion Mengesha
Connor Toups
John R. Rickford
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%