A survey on data‐efficient algorithms in big data era

Обзор алгоритмов с экономным использованием данных в эпоху больших данных
Amina Adadi
2021-01-26

data augmentationdata-efficient algorithmssmall-sample learningtransfer learningunsupervised learning
Abstract The leading approaches in Machine Learning are notoriously data-hungry. Unfortunately, many application domains do not have access to big data because acquiring data involves a process that is expensive or time-consuming. This has triggered a serious debate in both the industrial and academic communities calling for more data-efficient models that harness the power of artificial learners while achieving good results with less training data and in particular less human supervision. In light of this debate, this work investigates the issue of algorithms’ data hungriness. First, it surveys the issue from different perspectives. Then, it presents a comprehensive review of existing data-efficient methods and systematizes them into four categories. Specifically, the survey covers solution strategies that handle data-efficiency by (i) using non-supervised algorithms that are, by nature, more data-efficient, by (ii) creating artificially more data, by (iii) transferring knowledge from rich-data domains into poor-data domains, or by (iv) altering data-hungry algorithms to reduce their dependency upon the amount of samples, in a way they can perform well in small samples regime. Each strategy is extensively reviewed and discussed. In addition, the emphasis is put on how the four strategies interplay with each other in order to motivate exploration of more robust and data-efficient algorithms. Finally, the survey delineates the limitations, discusses research challenges, and suggests future opportunities to advance the research on data-efficiency in machine learning.
1
Altering data-hungry algorithms to reduce sample dependency enables those algorithms to perform well in small-sample regimes.
2
Creating artificial data (data augmentation) is a distinct strategy to increase effective training samples and improve performance with limited real data.
3
Machine learning methods are generally data-hungry and many application domains lack large labeled datasets due to expensive or time-consuming data acquisition.
4
The paper identifies limitations, research challenges, and future opportunities to advance data-efficiency in machine learning.
5
The survey emphasizes interactions among the four strategies, suggesting combined use may yield more robust and data-efficient algorithms.
6
The survey organizes data-efficient approaches into four main strategies: unsupervised algorithms, data augmentation (creating artificial data), transfer learning, and modifying data-hungry algorithms to work in small-sample regimes.
7
Transferring knowledge from rich-data domains to poor-data domains is a key approach to leverage existing labeled resources and improve performance where data is scarce.
8
Unsupervised algorithms are characterized as inherently more data-efficient and are a primary strategy for reducing reliance on labeled data.

data-efficient algorithms in machine learning

methods and strategies to reduce algorithms' data hungriness, enabling good performance with limited training data and reduced supervision (including unsupervised methods, data augmentation, transfer learning, and algorithmic adaptations)

Publication Details
Publication Date
2021-01-26
Journal
Publisher
ISSN
Cited by
332
Access Type
Author Information
Authors
Amina Adadi
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%