An evaluation of clustering and assembly strategies from Iso-Seq data in the absence of reference genomes in non-model animals

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Transcriptome assembly enables the recovery of expressed genes and isoforms, but the optimal strategy for reconstructing transcriptomes from long-read sequencing remains unresolved. In particular, establishing best practices for generating accurate gene models and selecting representative isoforms is essential for comparative genomics, since orthology inference typically requires only the longest isoform per gene model. Here, we systematically compare clustering and de novo assembly methods using PacBio Iso-Seq data from diverse invertebrate lineages with the goal of identifying the most optimal methodology for isoform selection in the absence of dedicated pipelines. We evaluate four approaches: IsoSeq3 (isoseq3 cluster), CD-HIT, RNA-Bloom2 and isONform, all benchmarked against short-read Trinity assemblies. Assembly quality was assessed using BUSCO completeness, short-read mapping rates, coding sequence recovery, longest isoform prediction, and SQANTI3 structural classification. Our results show that CD-HIT clustering at high similarity thresholds (≥99%) yields the most complete and coding-rich long-read transcriptomes, rivaling Trinity while avoiding its high redundancy. SQANTI3 classification further confirms that CD-HIT 99 recovers the highest number of full-splice-match transcripts among all methods. Consensus-based methods such as IsoSeq3 and isONform recover fewer single-copy orthologs (mirrored in a lower BUSCO score) and achieve lower mapping rates, while RNA-Bloom2 provides intermediate performance with reduced duplication. Together, these findings establish, to date, CD-HIT as a robust and practical strategy for transcriptome reconstruction from long-read data when genomic references are unavailable. This work provides practical guidance for deriving high-quality gene models and selecting representative isoforms for orthology inference in non-model species.

Article activity feed