Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences

Биологическая структура и функции возникают при масштабировании обучения без учителя до 250 миллионов белковых последовательностей
C. Lawrence Zitnick, Siddharth Goyal, Rob Fergus, Zeming Lin, Tom Sercu, Alexander Rives, Myle Ott, Jason Liu, Joshua Meier, Demi Guo, Jerry Ma
2019-04-29

mutational effect predictionprotein language modelingprotein sequence representationsremote homologyunsupervised learning
Abstract In the field of artificial intelligence, a combination of scale in data and model capacity enabled by un-supervised learning has led to major advances in representation learning and statistical generation. In the life sciences, the anticipated growth of sequencing promises unprecedented data on natural sequence diversity. Protein language modeling at the scale of evolution is a logical step toward predictive and generative artificial intelligence for biology. To this end we use unsupervised learning to train a deep contextual language model on 86 billion amino acids across 250 million protein sequences spanning evolutionary diversity. The resulting model contains information about biological properties in its representations. The representations are learned from sequence data alone. The learned representation space has a multi-scale organization reflecting structure from the level of biochemical properties of amino acids to remote homology of proteins. Information about secondary and tertiary structure is encoded in the representations and can be identified by linear projections. Representation learning produces features that generalize across a range of applications, enabling state-of-the-art supervised prediction of mutational effect and secondary structure, and improving state-of-the-art features for long-range contact prediction.
1
An unsupervised contextual protein language model was trained on 86 billion amino acids from 250 million evolutionarily diverse protein sequences.
2
Protein secondary- and tertiary-structure information is embedded in the learned representations and can be extracted using linear projections.
3
Sequence-only representations encode biological information organized across multiple scales, from amino-acid biochemical properties to remote protein homology.
4
The learned features enable state-of-the-art supervised prediction of mutational effects and secondary structure.
5
The representations improve state-of-the-art features for predicting long-range contacts between protein residues.

250 million evolutionarily diverse protein sequences and their learned protein-language-model representations

The emergence and encoding of multiscale biological structure and function, including biochemical properties, remote homology, protein structure, mutational effects, and long-range contacts

Publication Details
Publication Date
2019-04-29
Journal
Publisher
ISSN
Access Type
Author Information
Authors
C. Lawrence Zitnick
Siddharth Goyal
Rob Fergus
Zeming Lin
Tom Sercu
Alexander Rives
Myle Ott
Jason Liu
Joshua Meier
Demi Guo
Jerry Ma
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%