Generative genomics accurately predicts future experimental results
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Realizing AI’s promise to accelerate biomedical research requires AI models that are both accurate and sufficiently flexible to capture the diversity of real-life experiments. Here, we describe a generative genomics framework for AI-based experimental prediction that mirrors the process of designing and conducting an experiment in the lab or clinic. We created GEM-1 (Generate Expression Model-1), an AI system that effectively models the enormous range of bulk and single-cell gene expression experiments performed by scientists and benchmarked its performance across multiple biological axes. GEM-1’s prediction of future gene expression experiments–RNA-seq data deposited in public archives after our training data cutoff–yielded accuracy comparable to the best-possible performance estimated by comparing the results of matched lab experiments. Overall, our approach illustrates the transformative potential of generative genomics for applications ranging from predicting cellular perturbations in vitro to de novo generation of data from large clinical cohorts.