DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome
DNABERT: предварительно обученная двунаправленная модель представлений энкодера на основе Transformers для языка ДНК в геноме
2021-02-03
SCID: 54.1/tphf47m3
Discuss with AI
DNA language modelDNABERTgenomic DNA sequencesregulatory element predictiontranscription factor binding sites
Figures from the paper
Abstract (AI)
MOTIVATION: Deciphering the language of non-coding DNA is one of the fundamental problems in genome research. Gene regulatory code is highly complex due to the existence of polysemy and distant semantic relationship, which previous informatics methods often fail to capture especially in data-scarce scenarios. RESULTS: To address this challenge, we developed a novel pre-trained bidirectional encoder representation, named DNABERT, to capture global and transferrable understanding of genomic DNA sequences based on up and downstream nucleotide contexts. We compared DNABERT to the most widely used programs for genome-wide regulatory elements prediction and demonstrate its ease of use, accuracy and efficiency. We show that the single pre-trained transformers model can simultaneously achieve state-of-the-art performance on prediction of promoters, splice sites and transcription factor binding sites, after easy fine-tuning using small task-specific labeled data. Further, DNABERT enables direct visualization of nucleotide-level importance and semantic relationship within input sequences for better interpretability and accurate identification of conserved sequence motifs and functional genetic variant candidates. Finally, we demonstrate that pre-trained DNABERT with human genome can even be readily applied to other organisms with exceptional performance. We anticipate that the pre-trained DNABERT model can be fined tuned to many other sequence analyses tasks. AVAILABILITY AND IMPLEMENTATION: The source code, pretrained and finetuned model for DNABERT are available at GitHub (https://github.com/jerryji1993/DNABERT). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Key Findings
1
A DNABERT model pretrained on the human genome transfers readily to other organisms while maintaining exceptional performance.
2
A single pretrained DNABERT model achieves state-of-the-art performance for promoter, splice-site, and transcription-factor-binding-site prediction after fine-tuning with small labeled datasets.
3
Compared with widely used genome-wide regulatory-element prediction programs, DNABERT provides improved ease of use, accuracy, and efficiency.
4
DNABERT is a bidirectional Transformer pretrained on genomic DNA contexts to capture transferable, global representations of nucleotide sequences.
5
DNABERT supports nucleotide-level importance visualization and semantic-relationship analysis, enabling interpretable motif discovery and identification of candidate functional genetic variants.
Research Object
DNABERT — a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model applied to DNA sequences/genome
Research Subject
the sequence-context representations and predictive performance of regulatory functions, including promoters, splice sites, transcription factor binding sites, conserved motifs, and functional genetic variants
Publication Details
Publication Date
2021-02-03
Journal
Publisher
ISSN
Cited by
1345
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai2
Cited by6
The language of proteins: NLP, machine learning & protein sequences2021
DNA language models are powerful predictors of genome-wide variant effects2023
DNA language model GROVER learns sequence context in the human genome2024
GENA-LM: a family of open-source foundational DNA language models for long sequences2024
Species-aware DNA language models capture regulatory elements and their evolution2024
Cross-species modeling of plant genomes at single-nucleotide resolution using a pretrained DNA language model2025