BEND: Benchmarking DNA Language Models on biologically meaningful tasks

BEND: бенчмаркинг языковых моделей ДНК на биологически значимых задачах
Frederikke Isa Marin, Felix Teufel, Marc Horlacher, Dennis Madsen, Dennis Pultz, Ole Winther, Wouter Boomsma
2023-11-21

BEND benchmarkDNA language modelsbiologically meaningful tasksgenome annotationlong-range genomic features
The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND.
1
BEND introduces a benchmark for DNA language models comprising realistic, biologically meaningful downstream tasks defined on the human genome.
2
BEND is publicly available as an open-source benchmark at the provided GitHub repository.
3
Current DNA language model embeddings capture only limited information about long-range genomic features.
4
Embeddings from current DNA language models approach expert-method performance on some benchmark tasks.
5
The benchmark addresses limitations of prior evaluations by incorporating genome-annotation challenges involving data length, scale, and sparsity.

human genomic DNA sequences and their encoded functional, non-coding, and regulatory elements

the ability of DNA language model embeddings to capture biologically meaningful genomic annotations, including long-range features, across realistic downstream tasks

Publication Details
Publication Date
2023-11-21
Journal
Publisher
ISSN
Cited by
30
Access Type
Author Information
Authors
Frederikke Isa Marin
Felix Teufel
Marc Horlacher
Dennis Madsen
Dennis Pultz
Ole Winther
Wouter Boomsma
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%