Length-dependent prediction of protein intrinsic disorder

Предсказание внутренней неупорядоченности белков с учетом длины
Predrag Radivojac, Kang Peng, Slobodan Vučetić, A. Keith Dunker, Zoran Obradović
2006-04-17

10-fold cross-validationVSL2 predictorsintrinsic protein disorderlength-dependent predictionshort disordered regions
BACKGROUND: Due to the functional importance of intrinsically disordered proteins or protein regions, prediction of intrinsic protein disorder from amino acid sequence has become an area of active research as witnessed in the 6th experiment on Critical Assessment of Techniques for Protein Structure Prediction (CASP6). Since the initial work by Romero et al. (Identifying disordered regions in proteins from amino acid sequences, IEEE Int. Conf. Neural Netw., 1997), our group has developed several predictors optimized for long disordered regions (>30 residues) with prediction accuracy exceeding 85%. However, these predictors are less successful on short disordered regions (< or =30 residues). A probable cause is a length-dependent amino acid compositions and sequence properties of disordered regions. RESULTS: We proposed two new predictor models, VSL2-M1 and VSL2-M2, to address this length-dependency problem in prediction of intrinsic protein disorder. These two predictors are similar to the original VSL1 predictor used in the CASP6 experiment. In both models, two specialized predictors were first built and optimized for short (< or = 30 residues) and long disordered regions (>30 residues), respectively. A meta predictor was then trained to integrate the specialized predictors into the final predictor model. As the 10-fold cross-validation results showed, the VSL2 predictors achieved well-balanced prediction accuracies of 81% on both short and long disordered regions. Comparisons over the VSL2 training dataset via 10-fold cross-validation and a blind-test set of unrelated recent PDB chains indicated that VSL2 predictors were significantly more accurate than several existing predictors of intrinsic protein disorder. CONCLUSION: The VSL2 predictors are applicable to disordered regions of any length and can accurately identify the short disordered regions that are often misclassified by our previous disorder predictors. The success of the VSL2 predictors further confirmed the previously observed differences in amino acid compositions and sequence properties between short and long disordered regions, and justified our approaches for modelling short and long disordered regions separately. The VSL2 predictors are freely accessible for non-commercial use at http://www.ist.temple.edu/disprot/predictorVSL2.php.
1
A meta-predictor integrates specialized short- and long-region predictors into the final VSL2 models.
2
Ten-fold cross-validation showed balanced prediction accuracy of 81% for both short and long disordered regions.
3
The results support distinct amino acid compositions and sequence properties for short versus long disordered regions, enabling improved detection of short regions.
4
VSL2 predictors significantly outperformed several existing intrinsic-disorder predictors on cross-validation and an independent blind-test set of unrelated PDB chains.
5
VSL2-M1 and VSL2-M2 address length-dependent intrinsic disorder prediction by separately modeling short (≤30 residues) and long (>30 residues) disordered regions.

Intrinsically disordered protein regions of short (≤30 residues) and long (>30 residues) lengths

Length-dependent prediction of intrinsic protein disorder from amino acid sequences, including the comparative accuracy of specialized and integrated VSL2 predictor models for short and long disordered regions

Publication Details
Publication Date
2006-04-17
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Predrag Radivojac
Kang Peng
Slobodan Vučetić
A. Keith Dunker
Zoran Obradović
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%