Augmented Datasheets for Speech Datasets and Ethical Decision-Making

Расширенные datasheet'ы для речевых наборов данных и этического принятия решений
Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, William Thong, Dora Zhao, Jerone T. A. Andrews, R. D. Bourke, Alice Xiang, Allison Koenecke
2023-06-12

augmented datasheet for speech datasetsdata-subject protectionethical decision-makingspeech data documentationspeech datasets
Speech datasets are crucial for training Speech Language Technologies (SLT); however, the lack of diversity of the underlying training data can lead to serious limitations in building equitable and robust SLT products, especially along dimensions of language, accent, dialect, variety, and speech impairment—and the intersectionality of speech features with socioeconomic and demographic features. Furthermore, there is often a lack of oversight on the underlying training data—commonly built on massive web-crawling and/or publicly available speech—with regard to the ethics of such data collection. To encourage standardized documentation of such speech data components, we introduce an augmented datasheet for speech datasets1, which can be used in addition to “Datasheets for Datasets” [78]. We then exemplify the importance of each question in our augmented datasheet based on in-depth literature reviews of speech data used in domains such as machine learning, linguistics, and health. Finally, we encourage practitioners—ranging from dataset creators to researchers—to use our augmented datasheet to better define the scope, properties, and limits of speech datasets, while also encouraging consideration of data-subject protection and user community empowerment. Ethical dataset creation is not a one-size-fits-all process, but dataset creators can use our augmented datasheet to reflexively consider the social context of related SLT applications and data sources in order to foster more inclusive SLT products downstream.
1
Ethical oversight is frequently missing for speech training data, especially when sourced from large-scale web-crawling or public speech corpora.
2
Practitioners are encouraged to use the augmented datasheet to define dataset scope, properties, limits, and to consider data-subject protection and community empowerment for more inclusive SLT.
3
Speech datasets often lack diversity across language, accent, dialect, variety, and speech impairment, harming equitable and robust SLT products.
4
The augmented datasheet is motivated and exemplified through in-depth literature reviews across machine learning, linguistics, and health domains to show importance of each question.
5
The authors introduce an augmented datasheet for speech datasets to supplement existing 'Datasheets for Datasets' and standardize documentation of speech data components.

Speech datasets used to train Speech Language Technologies (SLT)

Documentation and ethical decision-making about dataset scope, composition, diversity (language, accent, dialect, impairment), collection practices, data-subject protection, and limits via an augmented datasheet to promote inclusive and responsible SLT

Publication Details
Publication Date
2023-06-12
Journal
Publisher
ISSN
Cited by
26
Access Type
Author Information
Authors
Orestis Papakyriakopoulos
Anna Seo Gyeong Choi
William Thong
Dora Zhao
Jerone T. A. Andrews
R. D. Bourke
Alice Xiang
Allison Koenecke
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%