epiTCR: a highly sensitive predictor for TCR–peptide binding

epiTCR: высокочувствительный предиктор связывания TCR с пептидами
My-Diem Nguyen Pham, Thanh-Nhan Nguyen, Le Son Tran, Que-Tran Bui Nguyen, Thien‐Phuc Hoang Nguyen, Thi Mong Quynh Pham, Hoai‐Nghia Nguyen, Hoa Giang, Minh‐Duy Phan, Vy Nguyen
2023-04-24

BLOSUM62 encodingRandom ForestTCR CDR3β sequencesTCR–peptide bindingneoantigen prediction
MOTIVATION: Predicting the binding between T-cell receptor (TCR) and peptide presented by human leucocyte antigen molecule is a highly challenging task and a key bottleneck in the development of immunotherapy. Existing prediction tools, despite exhibiting good performance on the datasets they were built with, suffer from low true positive rates when used to predict epitopes capable of eliciting T-cell responses in patients. Therefore, an improved tool for TCR-peptide prediction built upon a large dataset combining existing publicly available data is still needed. RESULTS: We collected data from five public databases (IEDB, TBAdb, VDJdb, McPAS-TCR, and 10X) to form a dataset of >3 million TCR-peptide pairs, 3.27% of which were binding interactions. We proposed epiTCR, a Random Forest-based method dedicated to predicting the TCR-peptide interactions. epiTCR used simple input of TCR CDR3β sequences and antigen sequences, which are encoded by flattened BLOSUM62. epiTCR performed with area under the curve (0.98) and higher sensitivity (0.94) than other existing tools (NetTCR, Imrex, ATM-TCR, and pMTnet), while maintaining comparable prediction specificity (0.9). We identified seven epitopes that contributed to 98.67% of false positives predicted by epiTCR and exerted similar effects on other tools. We also demonstrated a considerable influence of peptide sequences on prediction, highlighting the need for more diverse peptides in a more balanced dataset. In conclusion, epiTCR is among the most well-performing tools, thanks to the use of combined data from public sources and its use will contribute to the quest in identifying neoantigens for precision cancer immunotherapy. AVAILABILITY AND IMPLEMENTATION: epiTCR is available on GitHub (https://github.com/ddiem-ri-4D/epiTCR).
1
A dataset combining five public databases contains over 3 million TCR–peptide pairs, with 3.27% representing binding interactions.
2
Peptide sequence composition strongly influenced prediction performance, emphasizing the need for more diverse peptides and balanced datasets.
3
Seven epitopes accounted for 98.67% of epiTCR’s false positives and produced similar effects across other prediction tools.
4
The method maintained comparable prediction specificity of 0.90 relative to existing tools.
5
epiTCR achieved an area under the curve of 0.98 and sensitivity of 0.94, outperforming NetTCR, Imrex, ATM-TCR, and pMTnet in sensitivity.
6
epiTCR predicts TCR–peptide binding using a Random Forest model based on CDR3β and antigen sequences encoded with flattened BLOSUM62.

TCR–peptide binding interactions, represented by TCR CDR3β and antigen peptide sequences

Prediction performance and sequence-dependent determinants of TCR–peptide binding, particularly sensitivity, specificity, and false-positive behavior

Publication Details
Publication Date
2023-04-24
Journal
Publisher
ISSN
Cited by
68
Access Type
Author Information
Authors
My-Diem Nguyen Pham
Thanh-Nhan Nguyen
Le Son Tran
Que-Tran Bui Nguyen
Thien‐Phuc Hoang Nguyen
Thi Mong Quynh Pham
Hoai‐Nghia Nguyen
Hoa Giang
Minh‐Duy Phan
Vy Nguyen
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%