Enhancing TCR specificity predictions by combined pan- and peptide-specific training, loss-scaling, and sequence similarity integration
Повышение точности предсказания специфичности TCR за счет комбинирования панспецифического и пептид-специфического обучения, масштабирования функции потерь и интеграции сходства последовательностей
2023-12-20
SCID: 54.1/tpxpeqze
Discuss with AI
MHC class I peptidesNetTCR 2.2TCR specificity predictionpan-specific modelingpeptide-specific modeling
Figures from the paper
Abstract (AI)
Predicting the interaction between Major Histocompatibility Complex (MHC) class I-presented peptides and T-cell receptors (TCR) holds significant implications for vaccine development, cancer treatment, and autoimmune disease therapies. However, limited paired-chain TCR data, skewed towards well-studied epitopes, hampers the development of pan-specific machine-learning (ML) models. Leveraging a larger peptide-TCR dataset, we explore various alterations to the ML architectures and training strategies to address data imbalance. This leads to an overall improved performance, particularly for peptides with scant TCR data. However, challenges persist for unseen peptides, especially those distant from training examples. We demonstrate that such ML models can be used to detect potential outliers, which when removed from training, leads to augmented performance. Integrating pan-specific and peptide-specific models alongside with similarity-based predictions, further improves the overall performance, especially when a low false positive rate is desirable. In the context of the IMMREP22 benchmark, this modeling framework attained state-of-the-art performance. Moreover, combining these strategies results in acceptable predictive accuracy for peptides characterized with as little as 15 positive TCRs. This observation places great promise on rapidly expanding the peptide covering of the current models for predicting TCR specificity. The NetTCR 2.2 model incorporating these advances is available on GitHub (https://github.com/mnielLab/NetTCR-2.2) and as a web server at https://services.healthtech.dtu.dk/services/NetTCR-2.2/.
Key Findings
1
Combining pan-specific and peptide-specific training with loss scaling improves TCR specificity prediction, particularly for peptides with limited TCR data.
2
Integrating pan-specific, peptide-specific, and sequence-similarity predictions provides additional gains, especially when low false-positive rates are required.
3
Model performance improves overall, but predictions remain challenging for unseen peptides that are distant from training examples.
4
The combined framework achieved state-of-the-art performance on the IMMREP22 benchmark and acceptable accuracy for peptides with as few as 15 positive TCRs.
5
The models can identify potential outlier samples; removing these outliers from training further enhances predictive performance.
Research Object
MHC class I-presented peptides and their paired-chain T-cell receptors (TCRs)
Research Subject
TCR–peptide interaction specificity and the predictive performance of pan-specific, peptide-specific, loss-scaled, and sequence-similarity-integrated machine-learning models, particularly for data-scarce and unseen peptides
Publication Details
Publication Date
2023-12-20
Journal
Publisher
ISSN
Cited by
42
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest