State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis

Современные расширенные NLP-трансформеры для прямого и одношагового ретросинтеза
Igor V. Tetko, Pavel Karpov, Ruud Van Deursen, Guillaume Godin
2020-11-04

SMILES augmentationTransformerUSPTO-50k datasetbeam searchretrosynthesis prediction
We investigated the effect of different training scenarios on predicting the (retro)synthesis of chemical compounds using text-like representation of chemical reactions (SMILES) and Natural Language Processing (NLP) neural network Transformer architecture. We showed that data augmentation, which is a powerful method used in image processing, eliminated the effect of data memorization by neural networks and improved their performance for prediction of new sequences. This effect was observed when augmentation was used simultaneously for input and the target data simultaneously. The top-5 accuracy was 84.8% for the prediction of the largest fragment (thus identifying principal transformation for classical retro-synthesis) for the USPTO-50k test dataset, and was achieved by a combination of SMILES augmentation and a beam search algorithm. The same approach provided significantly better results for the prediction of direct reactions from the single-step USPTO-MIT test set. Our model achieved 90.6% top-1 and 96.1% top-5 accuracy for its challenging mixed set and 97% top-5 accuracy for the USPTO-MIT separated set. It also significantly improved results for USPTO-full set single-step retrosynthesis for both top-1 and top-10 accuracies. The appearance frequency of the most abundantly generated SMILES was well correlated with the prediction outcome and can be used as a measure of the quality of reaction prediction.
1
Applying SMILES data augmentation simultaneously to input and target eliminates neural network memorization and improves prediction of new reaction sequences.
2
Augmented transformer models significantly improved single-step retrosynthesis performance on the USPTO-full set for both top-1 and top-10 accuracies.
3
Combination of SMILES augmentation and beam search achieved 84.8% top-5 accuracy for predicting the largest fragment on the USPTO-50k test dataset.
4
For single-step direct reaction prediction on USPTO-MIT, the augmented transformer achieved 90.6% top-1 and 97% top-5 accuracy on the separated set, and 96.1% top-5 on a challenging mixed set.
5
The frequency of the most abundantly generated SMILES correlates well with prediction correctness and can serve as a quality measure for reaction predictions.

Transformer-based NLP models for predicting single-step retrosynthesis and direct chemical reactions using SMILES representations

Effect of SMILES data augmentation and training/decoding strategies (including simultaneous input-target augmentation and beam search) on prediction accuracy (top-1/top-5/top-10) and generation-frequency as a quality measure for single-step retrosynthesis and direct reaction prediction

Publication Details
Publication Date
2020-11-04
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Igor V. Tetko
Pavel Karpov
Ruud Van Deursen
Guillaume Godin
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%