GENA-LM: A Family of Open-Source Foundational DNA Language Models for Long Sequences
GENA-LM: семейство фундаментальных языковых моделей ДНК с открытым исходным кодом для длинных последовательностей
2023-06-13
SCID: 54.1/jmqnf5pf
Discuss with AI
DNA language modelsGENA-LMRecurrent Memorygenomicslong-sequence modeling
Figures from the paper
Abstract (AI)
Recent advancements in genomics, propelled by artificial intelligence, have unlocked unprecedented capabilities in interpreting genomic sequences, mitigating the need for exhaustive experimental analysis of complex, intertwined molecular processes inherent in DNA function. A significant challenge, however, resides in accurately decoding genomic sequences, which inherently involves comprehending rich contextual information dispersed across thousands of nucleotides. To address this need, we introduce GENA-LM, a suite of transformer-based foundational DNA language models capable of handling input lengths up to 36,000 base pairs. Notably, integration of the newly-developed Recurrent Memory mechanism allows these models to process even larger DNA segments. We provide pre-trained versions of GENA-LM, demonstrating their capability for fine-tuning and addressing a spectrum of complex biological tasks with modest computational demands. While language models have already achieved significant breakthroughs in protein biology, GENA-LM showcases a similarly promising potential for reshaping the landscape of genomics and multi-omics data analysis. All models are publicly available on GitHub https://github.com/AIRI-Institute/GENA LM and HuggingFace https://huggingface.co/AIRI-Institute.
Key Findings
1
A newly developed Recurrent Memory mechanism enables GENA-LM models to process DNA segments longer than their standard input limit.
2
GENA-LM is a suite of transformer-based foundational DNA language models supporting input sequences up to 36,000 base pairs.
3
GENA-LM models and their pre-trained versions are publicly released through GitHub and Hugging Face.
4
The models demonstrate promising potential for genomic and multi-omics data analysis by capturing long-range contextual information in DNA sequences.
5
The pre-trained models can be fine-tuned for diverse complex biological tasks with modest computational requirements.
Research Object
long genomic DNA sequences
Research Subject
contextual sequence modeling and downstream biological task performance for sequences up to 36,000 base pairs and beyond
Publication Details
Publication Date
2023-06-13
Journal
Publisher
ISSN
Cited by
56
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest