An alignment-last approach enables rapid transcriptomic biomarker discovery in large cohorts

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Canonical transcriptomic analysis requires committing from the outset to a reference genome or transcriptome, which imposes a predefined feature set, usually annotated genes or isoforms. Alignment and annotation dilute the signal through feature-level aggregation, discard any sequence absent from the reference, and require reprocessing the entire dataset for each new question (mutations, fusions, transposable elements). Here, we introduce the alignment-last paradigm, in which the read becomes the unit of comparison across samples, and alignment is deferred to annotate only the relevant sequences. Querying the merome , a reference-free cohort k-mer index, with just a handful of reads (about 0.01% of a sample’s) reveals the cohort’s transcriptomic structure in bulk and single-cell data. At single-cell resolution, these reads outperform genes for cell classification and rediscover, without supervision, a transposable-element signature (VL30) of exhausted T cells. Finally, unsupervised read-level differential analysis recovers established lncRNA biomarkers; uncovers new prognostic transposable-element reads in adrenocortical carcinoma and sarcomas; and extracts signals even from reads that fail to align.

Article activity feed