The language of proteins: NLP, machine learning & protein sequences

Язык белков: обработка естественного языка, машинное обучение и белковые последовательности
Michal Linial, Dan Ofer, Nadav Brandes
2021-01-01

contextualized embeddingsmasked language modelingnatural language processingprotein sequencesself-supervised learning
Natural language processing (NLP) is a field of computer science concerned with automated text and language analysis. In recent years, following a series of breakthroughs in deep and machine learning, NLP methods have shown overwhelming progress. Here, we review the success, promise and pitfalls of applying NLP algorithms to the study of proteins. Proteins, which can be represented as strings of amino-acid letters, are a natural fit to many NLP methods. We explore the conceptual similarities and differences between proteins and language, and review a range of protein-related tasks amenable to machine learning. We present methods for encoding the information of proteins as text and analyzing it with NLP methods, reviewing classic concepts such as bag-of-words, k-mers/n-grams and text search, as well as modern techniques such as word embedding, contextualized embedding, deep learning and neural language models. In particular, we focus on recent innovations such as masked language modeling, self-supervised learning and attention-based models. Finally, we discuss trends and challenges in the intersection of NLP and protein research.
1
Applying NLP to proteins shows substantial promise but also involves methodological pitfalls and unresolved challenges at the intersection of NLP and protein research.
2
Protein amino-acid sequences can be represented as text, making them naturally amenable to many natural language processing methods.
3
Protein analysis can use methods ranging from bag-of-words, k-mers or n-grams, and text search to word embeddings, contextualized embeddings, deep learning, and neural language models.
4
Recent protein modeling innovations include masked language modeling, self-supervised learning, and attention-based architectures.
5
The review compares proteins and human language, identifying both conceptual similarities and important differences relevant to computational modeling.

protein sequences represented as strings of amino-acid letters

the applicability, capabilities, and limitations of NLP and machine-learning methods for analyzing protein sequences and protein-related tasks

Publication Details
Publication Date
2021-01-01
Journal
Publisher
ISSN
Cited by
392
Access Type
Author Information
Authors
Michal Linial
Dan Ofer
Nadav Brandes
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%