Performance comparison of TCR-pMHC prediction tools reveals a strong data dependency
Сравнение эффективности инструментов предсказания TCR–pMHC выявляет сильную зависимость от данных
2023-04-18
SCID: 54.1/qgqf23cp
Discuss with AI
TCR-pMHC binding predictiondata imbalancedataset generalizationdeep learning modelsmodel robustness
Figures from the paper
Abstract (AI)
The interaction of T-cell receptors with peptide-major histocompatibility complex molecules (TCR-pMHC) plays a crucial role in adaptive immune responses. Currently there are various models aiming at predicting TCR-pMHC binding, while a standard dataset and procedure to compare the performance of these approaches is still missing. In this work we provide a general method for data collection, preprocessing, splitting and generation of negative examples, as well as comprehensive datasets to compare TCR-pMHC prediction models. We collected, harmonized, and merged all the major publicly available TCR-pMHC binding data and compared the performance of five state-of-the-art deep learning models (TITAN, NetTCR-2.0, ERGO, DLpTCR and ImRex) using this data. Our performance evaluation focuses on two scenarios: 1) different splitting methods for generating training and testing data to assess model generalization and 2) different data versions that vary in size and peptide imbalance to assess model robustness. Our results indicate that the five contemporary models do not generalize to peptides that have not been in the training set. We can also show that model performance is strongly dependent on the data balance and size, which indicates a relatively low model robustness. These results suggest that TCR-pMHC binding prediction remains highly challenging and requires further high quality data and novel algorithmic approaches.
Key Findings
1
A comprehensive benchmark compares five deep-learning models—TITAN, NetTCR-2.0, ERGO, DLpTCR, and ImRex—on merged public binding data.
2
All five evaluated models fail to generalize effectively to peptides absent from their training data.
3
Model performance depends strongly on dataset size and peptide-class balance, indicating relatively low robustness.
4
TCR-pMHC binding prediction remains challenging and requires higher-quality data and new algorithmic approaches.
5
The study introduces a general workflow for collecting, harmonizing, preprocessing, splitting, and generating negatives for TCR-pMHC binding datasets.
Research Object
T-cell receptor–peptide-major histocompatibility complex (TCR-pMHC) binding prediction models
Research Subject
Model generalization and robustness as functions of data splitting, dataset size, and peptide-class balance
Publication Details
Publication Date
2023-04-18
Journal
Publisher
ISSN
Cited by
54
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest