Benchmarking DNA Foundation Models for Genomic and Genetic Tasks

Haonan Feng
Lang Wu
Bingxin Zhao
Chad Huff
Jianjun Zhang
Jia Wu
Lifeng Lin
Peng Wei
Chong Wu

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

The rapid evolution of DNA foundation models promises to revolutionize genomics, yet comprehensive evaluations are lacking. Here, we present a comprehensive, unbiased benchmark of five models (DNABERT-2, Nucleotide Transformer V2, HyenaDNA, Caduceus-Ph, and GROVER) across diverse genomic and genetic tasks including sequence classification, gene expression prediction, variant effect quantification, and topologically associating domain (TAD) region recognition, using zero-shot embeddings. Our analysis reveals that mean token embedding consistently and significantly improves sequence classification performance, outperforming other pooling strategies. Model performance varies among tasks and datasets; while general purpose DNA foundation models showed competitive performance in pathogenic variant identification, they were less effective in predicting gene expression and identifying putative causal QTLs compared to specialized models. Our findings offer a framework for model selection, highlighting the impact of architecture, pre-training data, and embedding strategies on performance in genomic and genetic tasks.

Version published to 10.1101/2024.08.16.608288 on bioRxiv
Aug 18, 2024

NucleicBERT: Deciphering the language of nucleic acids by a large-language model

This article has 4 authors:
1. Utkarsh Upadhyay
2. Julian Herold
3. Markus Götz
4. Alexander Schug
This article has no evaluationsLatest version Sep 6, 2025
SNPBag: a foundation model for multitask genome-scale SNP analysis

This article has 7 authors:
1. Augix Xu
2. Yu Xu
3. Yiming Xing
4. Pengchao Luo
5. Jianbo Yang
6. Yinqi Bai
7. Kun Tang
This article has no evaluationsLatest version Sep 12, 2025
Predicting dynamic expression patterns in budding yeast with a fungal DNA language model

This article has 6 authors:
1. Kuan-Hao Chao
2. Majed Mohamed Magzoub
3. Emily Stoops
4. Sean Hackett
5. Johannes Linder
6. David R. Kelley
This article has no evaluationsLatest version Sep 21, 2025

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

NucleicBERT: Deciphering the language of nucleic acids by a large-language model

SNPBag: a foundation model for multitask genome-scale SNP analysis

Predicting dynamic expression patterns in budding yeast with a fungal DNA language model