GENA-LM: a family of open-source foundational DNA language models for long sequences
GENA-LM: семейство открытых фундаментальных языковых моделей ДНК для длинных последовательностей
2024-12-26
SCID: 54.1/jwvgvccw
Discuss with AI
DNA language modelsGENA-LMgenomic sequence analysisrecurrent memory mechanismtransformer-based models
Figures from the paper
Abstract (AI)
Recent advancements in genomics, propelled by artificial intelligence, have unlocked unprecedented capabilities in interpreting genomic sequences, mitigating the need for exhaustive experimental analysis of complex, intertwined molecular processes inherent in DNA function. A significant challenge, however, resides in accurately decoding genomic sequences, which inherently involves comprehending rich contextual information dispersed across thousands of nucleotides. To address this need, we introduce GENA language model (GENA-LM), a suite of transformer-based foundational DNA language models capable of handling input lengths up to 36 000 base pairs. Notably, integrating the newly developed recurrent memory mechanism allows these models to process even larger DNA segments. We provide pre-trained versions of GENA-LM, including multispecies and taxon-specific models, demonstrating their capability for fine-tuning and addressing a spectrum of complex biological tasks with modest computational demands. While language models have already achieved significant breakthroughs in protein biology, GENA-LM showcases a similarly promising potential for reshaping the landscape of genomics and multi-omics data analysis. All models are publicly available on GitHub (https://github.com/AIRI-Institute/GENA_LM) and on HuggingFace (https://huggingface.co/AIRI-Institute). In addition, we provide a web service (https://dnalm.airi.net/) allowing user-friendly DNA annotation with GENA-LM models.
Key Findings
1
A recurrent memory mechanism enables GENA-LM models to process DNA segments longer than their standard input limit.
2
GENA-LM is a suite of open-source transformer-based foundational DNA language models supporting input sequences up to 36,000 base pairs.
3
GENA-LM models can address genomic and multi-omics analysis tasks with modest computational requirements.
4
Pretrained models, source code, and a web-based DNA annotation service are publicly available through GitHub, Hugging Face, and dnalm.airi.net.
5
The suite includes multispecies and taxon-specific pretrained models designed for fine-tuning across diverse complex biological tasks.
Research Object
GENA-LM family of transformer-based foundational DNA language models for long genomic sequences (up to 36,000 bp) including multispecies and taxon-specific variants
Research Subject
contextual sequence representation and modeling capabilities over up to 36,000 base pairs, including scalability to larger DNA segments and transferability to biological tasks
Publication Details
Publication Date
2024-12-26
Journal
Publisher
ISSN
Cited by
82
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai5
DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome2021
Longformer: The Long-Document Transformer2020
ETC: Encoding Long and Structured Inputs in Transformers2020
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
Proceedings of the 24th international conference on Machine learning2007