Investigating the determinants of performance in machine learning for protein fitness prediction

Исследование факторов, определяющих эффективность машинного обучения при прогнозировании приспособленности белков
Mahakaran Sandhu, Adam C. Mater, Dana S. Matthews, Matthew A. Spence, Artem Lenskiy, Colin J. Jackson
2025-07-21

epistasis and ruggednessfitness landscapespositional extrapolationprotein fitness predictionsequence-fitness prediction
Machine learning (ML) has revolutionized protein biology, solving long-standing problems in protein folding, scaffold generation, and function design tasks. A range of architectures have shown success on supervised protein fitness prediction tasks. Nevertheless, in the absence of rational approaches for evaluating which architectures are optimal for specific datasets and engineering tasks, architecture choice remains challenging. Here, we propose a framework for investigating the determinants of success for a range of ML architectures. Using simulated (the NK model) and empirical fitness landscapes, we measure sequence-fitness prediction along six key performance metrics: interpolation within the training domain, extrapolation outside the training domain, robustness to increasing epistasis/ruggedness, ability to perform positional extrapolation, robustness to sparse training data, and sensitivity to sequence length. We show that architectural differences between algorithms consistently affect performance against these metrics across both experimental and theoretical landscapes. Moreover, landscape ruggedness emerges as a primary determinant of the accuracy of sequence-fitness prediction. Our methodology and results provide a rational strategy for experimental data sampling, model selection, and evaluation rooted in fitness landscape theory-one that we hope will advance sequence-fitness prediction accuracy, with implications for protein engineering and variant functional prediction.
1
Architectural differences consistently influence performance across both simulated NK-model and empirical protein fitness landscapes.
2
Fitness-landscape ruggedness is identified as a primary determinant of sequence–fitness prediction accuracy.
3
Performance is evaluated across interpolation, extrapolation, epistasis robustness, positional extrapolation, sparse-data robustness, and sequence-length sensitivity.
4
The framework supports rational experimental data sampling, model selection, and evaluation for protein engineering and variant-function prediction.
5
The paper introduces a framework for identifying why different machine-learning architectures succeed or fail in protein sequence–fitness prediction.

machine learning architectures for protein sequence-fitness prediction on simulated and empirical fitness landscapes

determinants of prediction performance, including interpolation and extrapolation, robustness to landscape ruggedness and sparse data, positional extrapolation, and sensitivity to sequence length

Publication Details
Publication Date
2025-07-21
Journal
Publisher
ISSN
Cited by
6
Access Type
Author Information
Authors
Mahakaran Sandhu
Adam C. Mater
Dana S. Matthews
Matthew A. Spence
Artem Lenskiy
Colin J. Jackson
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%