Machine Learning Methods for Small Data Challenges in Molecular Science

Методы машинного обучения для решения задач, связанных с малыми объемами данных в молекулярной науке
Guo‐Wei Wei, Bozheng Dou, Zailiang Zhu, Ekaterina Merkurjev, Ke Lü, Long Chen, Jiang Jian, Yueying Zhu, Jie Liu, Bengong Zhang
2023-06-29

data augmentationdeep learningmachine learningmolecular sciencesmall data challenges
Small data are often used in scientific and engineering research due to the presence of various constraints, such as time, cost, ethics, privacy, security, and technical limitations in data acquisition. However, big data have been the focus for the past decade, small data and their challenges have received little attention, even though they are technically more severe in machine learning (ML) and deep learning (DL) studies. Overall, the small data challenge is often compounded by issues, such as data diversity, imputation, noise, imbalance, and high-dimensionality. Fortunately, the current big data era is characterized by technological breakthroughs in ML, DL, and artificial intelligence (AI), which enable data-driven scientific discovery, and many advanced ML and DL technologies developed for big data have inadvertently provided solutions for small data problems. As a result, significant progress has been made in ML and DL for small data challenges in the past decade. In this review, we summarize and analyze several emerging potential solutions to small data challenges in molecular science, including chemical and biological sciences. We review both basic machine learning algorithms, such as linear regression, logistic regression (LR), k -nearest neighbor (KNN), support vector machine (SVM), kernel learning (KL), random forest (RF), and gradient boosting trees (GBT), and more advanced techniques, including artificial neural network (ANN), convolutional neural network (CNN), U-Net, graph neural network (GNN), Generative Adversarial Network (GAN), long short-term memory (LSTM), autoencoder, transformer, transfer learning, active learning, graph-based semi-supervised learning, combining deep learning with traditional machine learning, and physical model-based data augmentation. We also briefly discuss the latest advances in these methods. Finally, we conclude the survey with a discussion of promising trends in small data challenges in molecular science.
1
Combining machine learning with domain knowledge and physical models is identified as a promising direction for future molecular-science applications.
2
Methods originally developed for big-data machine learning and deep learning have enabled substantial progress in addressing small-data problems.
3
Small datasets create severe machine-learning challenges involving limited diversity, missing values, noise, class imbalance, and high dimensionality.
4
Small-data constraints in molecular science arise from time, cost, ethical, privacy, security, and technical limitations in data acquisition.
5
The review analyzes classical algorithms and advanced approaches, including transfer learning, active learning, semi-supervised learning, generative models, and physics-based data augmentation.

small-data problems in molecular science, including chemical and biological sciences

machine-learning and deep-learning solutions for handling data diversity, imputation, noise, imbalance, and high dimensionality in molecular-science small-data applications

Publication Details
Publication Date
2023-06-29
Journal
Publisher
ISSN
Cited by
497
Access Type
Author Information
Authors
Guo‐Wei Wei
Bozheng Dou
Zailiang Zhu
Ekaterina Merkurjev
Ke Lü
Long Chen
Jiang Jian
Yueying Zhu
Jie Liu
Bengong Zhang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%