Detecting Random Mutations in 16S rRNA Sequences
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Motivation
High-throughput sequencing technologies have driven rapid growth of biological sequence databases. Public repositories must therefore rely on automated computational heuristics to screen submitted sequences for errors and low quality. For example, SILVA SSU Ref, which exploits the conserved nature of 16S and 18S rRNA sequences, applies strict algorithmic quality controls yet still accepts sequences with up to 30% of their nucleotides deviating from any previously accepted sequence. This permissiveness creates opportunities for the admission of modified sequences, such as biologically plausible sequences generated by DNA foundation models. The vulnerability of public sequence databases to becoming polluted or poisoned with modified sequences necessitates the development of methods to detect such sequences.
Results
We present the first investigation, to our knowledge, of the detectability of modified sequences. We consider simple computationally-modified 16S rRNA sequences that pass the quality control inclusion criteria of the SILVA database, which we generate via random substitutions. We present classifiers that can distinguish such modified sequences from natural 16S rRNA using conserved motifs. Our best classifier achieves over 90% sensitivity and specificity on our testing set when using a 5% artificial mutation rate. One feature used in our classifiers, gapped k-mers constructed from universally conserved nucleotides, was conserved across all three domains of life despite relying on exact matches to patterns found in E. coli , advancing our understanding of conserved grammatical structure in small subunit rRNA sequences.
Availability and Implementation
Our source code is available at https://github.com/rainhaworth/16S-Mutation-Classifiers .