Pretraining Enhances Megabase-Scale Gene Expression Prediction with GeneUnet

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Predicting gene expression from DNA sequence across diverse genomic tracks is essential for understanding gene regulation and interpreting non-coding variants. Existing supervised methods are limited to few species and fail to exploit conserved regulatory mechanisms, while DNA foundation models capture cross-species information but remain constrained to kilobase-scale contexts insufficient for this task. Here we introduce GB.GeneUnet, an 837M-parameter transformer-based U-Net pretrained on 6 trillion tokens from multi-species genomes in OpenGenome2, extending genomic context to 1 Mb with up to 100× inference speedup over GeneMoE, a preliminary MoE transformer baseline of similar model size pretrained on the same data. Fine-tuned for gene expression prediction, GB.GeneUnet achieves state-of-the-art performance on the Borzoi benchmark at 524 kb context, and attains performance comparable to AlphaGenome at 1 Mb context while requiring a lighter fine-tuning procedure. Together, these results establish a scalable framework linking multi-species pretraining to ultra-long-context gene expression modeling.

Article activity feed