Text Data Augmentation for Deep Learning

Аугментация текстовых данных для глубокого обучения
Connor Shorten, Taghi M. Khoshgoftaar, Borko Furht
2021-07-19

consistency regularizationcounterfactual examplesgeneralization testingnatural language processingtext data augmentation
Natural Language Processing (NLP) is one of the most captivating applications of Deep Learning. In this survey, we consider how the Data Augmentation training strategy can aid in its development. We begin with the major motifs of Data Augmentation summarized into strengthening local decision boundaries, brute force training, causality and counterfactual examples, and the distinction between meaning and form. We follow these motifs with a concrete list of augmentation frameworks that have been developed for text data. Deep Learning generally struggles with the measurement of generalization and characterization of overfitting. We highlight studies that cover how augmentations can construct test sets for generalization. NLP is at an early stage in applying Data Augmentation compared to Computer Vision. We highlight the key differences and promising ideas that have yet to be tested in NLP. For the sake of practical implementation, we describe tools that facilitate Data Augmentation such as the use of consistency regularization, controllers, and offline and online augmentation pipelines, to preview a few. Finally, we discuss interesting topics around Data Augmentation in NLP such as task-specific augmentations, the use of prior knowledge in self-supervised learning versus Data Augmentation, intersections with transfer and multi-task learning, and ideas for AI-GAs (AI-Generating Algorithms). We hope this paper inspires further research interest in Text Data Augmentation.
1
Data augmentation can construct test sets that assess generalization, addressing deep learning’s difficulty in measuring generalization and characterizing overfitting.
2
It catalogs concrete augmentation frameworks developed for textual data and discusses their practical implementation through consistency regularization, controllers, and offline or online pipelines.
3
NLP applies data augmentation less maturely than computer vision, leaving major opportunities to adapt promising computer-vision ideas and develop untested NLP methods.
4
The survey identifies open research directions including task-specific augmentation, prior knowledge in self-supervised learning, transfer and multi-task learning, and AI-generating algorithms.
5
The survey organizes text data augmentation around four motifs: strengthening local decision boundaries, brute-force training, causality and counterfactuals, and separating meaning from form.

text data in Natural Language Processing (NLP)

the principles, frameworks, and effects of data augmentation on NLP model training, generalization, and overfitting

Publication Details
Publication Date
2021-07-19
Journal
Publisher
ISSN
Cited by
1707
Access Type
Author Information
Authors
Connor Shorten
Taghi M. Khoshgoftaar
Borko Furht
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%