TEscape: Defining the human transposable element transcriptome using multiplatform long-read sequencing

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Transposable elements (TEs) not only account for half of the human genome sequence but also generate transcripts that contribute to transcriptomic diversity. Yet, their repetitive nature has hindered accurate quantification of the full TE-derived transcriptome, a challenge that long-read sequencing can overcome. Here, we combined multiplexed arrays isoform sequencing (MAS-ISO-seq) with a dedicated computational framework (TEscape) to perform an in-depth annotation of the human TE transcriptome. To capture the breadth of human transcriptome diversity, we profiled six representative cell types spanning three distinct biological contexts, including metabolism with, primary patient-derived adipogenic cells at two differentiation stages, and iPSC derived hepatic progenitor cells; the nervous system with iPSC-derived neurons, neural progenitor cells (NPCs), and pluripotency using induced pluripotent stem cells (iPSCs). Together, these datasets yielded over 235 million full-length long reads. First, to assess data coverage and transcriptome depth, we quantified protein-coding gene expression, detecting 14,312 genes (73.6% of all annotated protein-coding genes), which is a level consistent with deep and comprehensive transcriptome representation. Second, focusing on TE-derived transcripts, we identified >83,000 previously unannotated isoforms, the vast majority (84%) originating from a complex combination of multi-TEs. We also identified solo TEs, which are predominantly from LINE1 (14%). We confirmed that TE-transcripts are able to be exemplified by signatures detected in Liver Hepatocellular Carcinoma (LICH). Together, MAS-ISO-seq and TEscape establish the first long-read-based, high-resolution atlas of transcribed human TEs, providing a foundational resource for integrative transcriptome analyses and for investigating TE expression and regulation in health and disease.

Article activity feed