Improving machine-learning models in materials science through large datasets

Совершенствование моделей машинного обучения в материаловедении с помощью больших наборов данных
Miguel A. L. Marques, Jonathan Schmidt, Silvana Botti, Hai‐Chen Wang, A. Romero, Tiago F. T. Cerqueira, Antoine Loew, Fabian Jäger
2024-09-25

Alexandria databasecrystal-graph neural networksdensity-functional theory calculationsmachine-learning interatomic potentialsmaterials science datasets
The accuracy of a machine learning model is limited by the quality and quantity of the data available for its training and validation. This problem is particularly challenging in materials science, where large, high-quality, and consistent datasets are scarce. Here we present alexandria , an open database of more than 5 million density-functional theory calculations for periodic three-, two-, and one-dimensional compounds. We use this data to train machine learning models to reproduce seven different properties using both composition-based models and crystal-graph neural networks. In the majority of cases, the error of the models decreases monotonically with the training data, although some graph networks seem to saturate for large training set sizes. Differences in the training can be correlated with the statistical distribution of the different properties. We also observe that graph-networks, that have access to detailed geometrical information, yield in general more accurate models than simple composition-based methods. Finally, we assess several universal machine learning interatomic potentials. Crystal geometries optimised with these force fields are very high quality, but unfortunately the accuracy of the energies is still lacking. Furthermore, we observe some instabilities for regions of chemical space that are undersampled in the training sets used for these models. This study highlights the potential of large-scale, high-quality datasets to improve machine learning models in materials science.
1
Alexandria provides an open database containing more than 5 million density-functional theory calculations for periodic three-, two-, and one-dimensional compounds.
2
Crystal-graph neural networks generally outperform composition-based models, consistent with their access to detailed geometric information.
3
Machine-learning models were trained to predict seven properties using composition-based methods and crystal-graph neural networks.
4
Model errors generally decrease monotonically as training data increases, although some graph networks saturate at large training-set sizes.
5
Universal machine-learning interatomic potentials produce high-quality optimized geometries, but energy accuracy remains insufficient and instabilities occur in undersampled chemical-space regions.

Machine-learning models trained on large datasets of periodic three-, two-, and one-dimensional compounds from density-functional theory calculations

How training-dataset size, data quality, property distributions, model architecture, and chemical-space sampling affect the accuracy, convergence, and stability of machine-learning predictions and interatomic potentials for materials properties

Publication Details
Publication Date
2024-09-25
Journal
Publisher
ISSN
Cited by
112
Access Type
Author Information
Authors
Miguel A. L. Marques
Jonathan Schmidt
Silvana Botti
Hai‐Chen Wang
A. Romero
Tiago F. T. Cerqueira
Antoine Loew
Fabian Jäger
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%