Better data for better predictions: data curation improves deep learning for sgRNA/Cas9 prediction

Tyler S. Browne
David R. Edgell
Gregory B. Gloor

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

The Cas9 enzyme along with a single guide RNA molecule is a modular tool for genetic engineering and has shown effectiveness as a species-specific anti-microbial. The ability to accurately predict on-target cleavage is critical as activity varies by target. Using the sgRNA nucleotide sequence and an activity score, predictive models have been developed with the best performance resulting from deep learning architectures. Prior work has emphasized robust and novel architectures to improve predictive performance. Here, we explore the impact of a data-centric approach through optimization of the input target site adjacent nucleotide sequence length and the use of data filtering for read counts in the control conditions to improve input data utility. Using the existing crisprHAL architecture, we develop crisprHAL Tev, a best-in-class bacterial SpCas9 prediction model with performance that generalizes between related species and across data types. During this process, we also rebuild two prior E. coli Cas9 datasets, demonstrating the importance of data quality, and resulting in the production of an improved bacterial eSpCas9 prediction model. The crisprHAL models are available through GitHub ( https://github.com/tbrowne/crisprHAL ).

Version published to 10.1101/2025.06.24.661356 on bioRxiv
Jun 27, 2025

Decoupled Representation Learning Improves Generalization in CRISPR Off-Target Prediction

This article has 2 authors:
1. Nyla Bhargava
2. Aditya Goswami
This article has no evaluationsLatest version Jan 18, 2026
Deep Learning Approaches for Accurate RNA 3D Structure Prediction from Primary Sequences

This article has 1 author:
1. Nnaemeka Kingsley Ugwumba
This article has no evaluationsLatest version Jan 29, 2026
Convolutional Deep Learning Approach to identify DNA Sequences for Gene Prediction

This article has 2 authors:
1. Jesus Antonio Motta
2. Pedro David Gomez
This article has no evaluationsLatest version Jan 27, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Decoupled Representation Learning Improves Generalization in CRISPR Off-Target Prediction

Deep Learning Approaches for Accurate RNA 3D Structure Prediction from Primary Sequences

Convolutional Deep Learning Approach to identify DNA Sequences for Gene Prediction