Extended t -cores for the de novo identification of transposable elements and other inexact repeats from short read RNA-seq data

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Transcribed repeats represent a major challenge in the de novo assembly of transcriptomes from short RNA-seq reads. Young transposable elements (TEs) and more generally, inexact repeats, create dense and ambiguous regions in the assembly graph, preventing the correct assembly of transcripts. In this paper, we introduce a fully de novo method based on the discovery of dense regions in the compacted De Bruijn graph (DBG) to identify such repeats directly from short read RNA-seq data, without requiring a reference genome or repeat database. Our approach defines the extended t-cores , subgraphs of the DBG that capture the complex topology induced by expressed inexact repeats appearing in RNA-seq reads.

Independently of its interest for transcriptome assembly, the proposed method appears to be effective for the de novo identification of repeats in transcriptomes. After classifying cores using sequence-based motifs to distinguish simple repeats from potential TEs, we demonstrate its potential for the de novo discovery of TEs.

We validate the approach on a Mus musculus dataset, showing that extended t -cores correspond to known expressed TE families. We also illustrate its de novo discovery potential on a non-model species, Canis lupus familiaris , where the method was also able to recover known TEs.

Article activity feed