DNA language model GROVER learns sequence context in the human genome

Языковая модель ДНК GROVER изучает контекст последовательности в геноме человека
Pierre M. Joubert, Melissa Sanabria, Anna R. Poetsch, Jonas Hirsch
2024-07-23

GROVER DNA language modelbyte-pair encodinghuman genome sequence contextnext-k-mer predictionprotein–DNA binding
Abstract Deep-learning models that learn a sense of language on DNA have achieved a high level of performance on genome biological tasks. Genome sequences follow rules similar to natural language but are distinct in the absence of a concept of words. We established byte-pair encoding on the human genome and trained a foundation language model called GROVER (Genome Rules Obtained Via Extracted Representations) with the vocabulary selected via a custom task, next- k -mer prediction. The defined dictionary of tokens in the human genome carries best the information content for GROVER. Analysing learned representations, we observed that trained token embeddings primarily encode information related to frequency, sequence content and length. Some tokens are primarily localized in repeats, whereas the majority widely distribute over the genome. GROVER also learns context and lexical ambiguity. Average trained embeddings of genomic regions relate to functional genomics annotation and thus indicate learning of these structures purely from the contextual relationships of tokens. This highlights the extent of information content encoded by the sequence that can be grasped by GROVER. On fine-tuning tasks addressing genome biology with questions of genome element identification and protein–DNA binding, GROVER exceeds other models’ performance. GROVER learns sequence context, a sense for structure and language rules. Extracting this knowledge can be used to compose a grammar book for the code of life.
1
After fine-tuning, GROVER outperforms other models on genome-element identification and protein–DNA binding tasks, demonstrating learned sequence structure and language rules.
2
Average embeddings of genomic regions correlate with functional genomics annotations, indicating that functional structure can be learned from token-context relationships alone.
3
GROVER is a human-genome foundation language model trained with byte-pair encoding and a custom next-k-mer prediction task for vocabulary selection.
4
GROVER learns genomic context and lexical ambiguity; some tokens localize mainly in repeats, whereas most distribute broadly across the genome.
5
The learned genomic token dictionary captures information content effectively, while token embeddings primarily encode frequency, sequence content, and token length.

human genome DNA sequences

sequence context, language-like rules, and encoded genomic information learned from token representations, including implications for genome element identification and protein–DNA binding

Publication Details
Publication Date
2024-07-23
Journal
Publisher
ISSN
Cited by
101
Access Type
Author Information
Authors
Pierre M. Joubert
Melissa Sanabria
Anna R. Poetsch
Jonas Hirsch
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%