Species-aware DNA language models capture regulatory elements and their evolution

Видоспецифичные языковые модели ДНК выявляют регуляторные элементы и их эволюцию
Julien Gagneur, Alexander Karollus, Johannes Hingerl, Dennis Gankin, Martin Grosshauser, Kristian Klemon
2024-04-02

MPRA gene expression predictionevolutionary conservationregulatory elementsspecies-aware DNA language modelstranscription factor motifs
BACKGROUND: The rise of large-scale multi-species genome sequencing projects promises to shed new light on how genomes encode gene regulatory instructions. To this end, new algorithms are needed that can leverage conservation to capture regulatory elements while accounting for their evolution. RESULTS: Here, we introduce species-aware DNA language models, which we trained on more than 800 species spanning over 500 million years of evolution. Investigating their ability to predict masked nucleotides from context, we show that DNA language models distinguish transcription factor and RNA-binding protein motifs from background non-coding sequence. Owing to their flexibility, DNA language models capture conserved regulatory elements over much further evolutionary distances than sequence alignment would allow. Remarkably, DNA language models reconstruct motif instances bound in vivo better than unbound ones and account for the evolution of motif sequences and their positional constraints, showing that these models capture functional high-order sequence and evolutionary context. We further show that species-aware training yields improved sequence representations for endogenous and MPRA-based gene expression prediction, as well as motif discovery. CONCLUSIONS: Collectively, these results demonstrate that species-aware DNA language models are a powerful, flexible, and scalable tool to integrate information from large compendia of highly diverged genomes.
1
Species-aware DNA language models were trained on more than 800 species spanning over 500 million years of evolution.
2
Species-aware training improves sequence representations for endogenous and MPRA-based gene-expression prediction and motif discovery.
3
The models distinguish transcription-factor and RNA-binding-protein motifs from background non-coding DNA when predicting masked nucleotides.
4
The models reconstruct motifs bound in vivo better than unbound motifs and capture motif-sequence evolution and positional constraints.
5
Unlike sequence alignment, the models capture conserved regulatory elements across much larger evolutionary distances.

multi-species genomic DNA sequences, including conserved and evolving non-coding regulatory regions across more than 800 species

the capacity of species-aware DNA language models to capture transcription factor and RNA-binding protein motifs, their functional binding status, positional constraints, and evolutionary conservation

Publication Details
Publication Date
2024-04-02
Journal
Publisher
ISSN
Cited by
54
Access Type
Author Information
Authors
Julien Gagneur
Alexander Karollus
Johannes Hingerl
Dennis Gankin
Martin Grosshauser
Kristian Klemon
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%