The language of proteins: NLP, machine learning & protein sequences
Язык белков: обработка естественного языка, машинное обучение и белковые последовательности
2021-01-01
SCID: 54.1/b4p3ybng
Discuss with AI
contextualized embeddingsmasked language modelingnatural language processingprotein sequencesself-supervised learning
Figures from the paper
Abstract (AI)
Natural language processing (NLP) is a field of computer science concerned with automated text and language analysis. In recent years, following a series of breakthroughs in deep and machine learning, NLP methods have shown overwhelming progress. Here, we review the success, promise and pitfalls of applying NLP algorithms to the study of proteins. Proteins, which can be represented as strings of amino-acid letters, are a natural fit to many NLP methods. We explore the conceptual similarities and differences between proteins and language, and review a range of protein-related tasks amenable to machine learning. We present methods for encoding the information of proteins as text and analyzing it with NLP methods, reviewing classic concepts such as bag-of-words, k-mers/n-grams and text search, as well as modern techniques such as word embedding, contextualized embedding, deep learning and neural language models. In particular, we focus on recent innovations such as masked language modeling, self-supervised learning and attention-based models. Finally, we discuss trends and challenges in the intersection of NLP and protein research.
Key Findings
1
Applying NLP to proteins shows substantial promise but also involves methodological pitfalls and unresolved challenges at the intersection of NLP and protein research.
2
Protein amino-acid sequences can be represented as text, making them naturally amenable to many natural language processing methods.
3
Protein analysis can use methods ranging from bag-of-words, k-mers or n-grams, and text search to word embeddings, contextualized embeddings, deep learning, and neural language models.
4
Recent protein modeling innovations include masked language modeling, self-supervised learning, and attention-based architectures.
5
The review compares proteins and human language, identifying both conceptual similarities and important differences relevant to computational modeling.
Research Object
protein sequences represented as strings of amino-acid letters
Research Subject
the applicability, capabilities, and limitations of NLP and machine-learning methods for analyzing protein sequences and protein-related tasks
Publication Details
Publication Date
2021-01-01
Journal
Publisher
ISSN
Cited by
392
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai4
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome2021
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
Enriching Word Vectors with Subword Information2017