xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction

xDecoder раскрывает потенциал геномных фундаментальных моделей для предсказания индивидуальной экспрессии генов в режиме few-shot
Ruibang Luo, Yuanhua Huang, Shumin Li
2026-07-28

cross-locus transferfew-shot gene expression predictiongenomic language modelspersonalized DNA-RNA trainingxDecoder
Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to capture personal gene expression variations remains unresolved. We developed xDecoder, a unified decoding framework that utilizes gLMs and sequence-to-function (S2F) embeddings to learn how personal genetic variation shapes gene expression from paired genome-transcriptome data. Compared to the pretrained genomic models, xDecoder with personalized DNA-RNA training makes cross-individual prediction tractable for seen genes in a few-shot setting. However, zero-shot prediction at unseen loci remains unreliable and gene-dependent, revealing a cross-locus transfer bottleneck of current sequence models. Experiments incorporating individual-level chromatin accessibility suggested that regulatory-state information important for unseen-locus prediction is not fully captured by current DNA-only models. Overall, these results highlight the potential utility of the few-shot setting, the limitations of DNA-only models, and point toward multi-omic, variant-aware frameworks as a promising direction for building personalized regulatory models.
1
Adding individual-level chromatin accessibility indicates that DNA-only models incompletely capture regulatory-state information needed for unseen-locus prediction.
2
Personalized DNA-RNA training makes cross-individual gene-expression prediction tractable for previously observed genes in a few-shot setting.
3
The findings support few-shot personalized prediction while motivating multi-omic, variant-aware frameworks for improved regulatory modeling.
4
Zero-shot prediction at unseen genomic loci remains unreliable and gene-dependent, exposing a cross-locus transfer bottleneck in current sequence models.
5
xDecoder is a unified framework combining genomic language models and sequence-to-function embeddings to predict personal gene expression from paired genome-transcriptome data.

Personal gene expression shaped by individual genetic variation, modeled from paired genome–transcriptome data

Few-shot and zero-shot cross-individual and cross-locus prediction of gene expression, including the roles and limitations of DNA sequence, chromatin accessibility, and regulatory-state information

Publication Details
Publication Date
2026-07-28
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Ruibang Luo
Yuanhua Huang
Shumin Li
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%