Computing tumor specificity of cancer antigen targets by k-mer indexing of healthy tissue transcriptomes
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Individualized cancer immunotherapies rely on tumor-specific T-cell antigens, often predicted from somatic mutations as neoantigens. For tumors with low mutational burden, mRNA transcript variants, including gene fusions and novel splice junctions, can serve as important alternative targets. A main challenge in their identification from tumor RNA-seq is to confirm that their expression is tumor-restricted. Although large public collections of healthy-tissue RNA-seq exist, verifying tumor-specific expression requires computationally expensive re-analysis of these data for every novel candidate. To address this, we benchmarked nine k-mer indexing algorithms and developed k4neo, which leverages k-mer indexing of raw RNA-seq reads to compute the tumor specificity of any transcript variant. This mapping-free and transcript-class agnostic approach screens any candidate sequence against 18,960 samples across 51 healthy tissue types. We confirmed k4neo’s detection accuracy with qRT-PCR and showed that k4neo accurately classifies somatic and germline variants, gene fusions, and isoforms by tumor specificity. Applied to nine tumor cohorts, it nominated a median of 4-80 tumor-specific splice junctions per patient, including recurrent, long-read-confirmed novel antigen candidates. Together, k4neo enables efficient access to large-scale sequencing cohorts and accurately computes tumor specificity for any input transcript sequence, thereby expanding the repertoire of individual and shared cancer antigen targets.