Combining transcriptomic resolutions and machine learning strategies uncovers new OXPHOS genes in Caenorhabditis elegans

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Assigning functions to genes remains a major bottleneck in biology, as many genes remain uncharacterized despite the availability of complete genome sequences. Oxidative phosphorylation (OXPHOS), the primary source of ATP in eukaryotes, exemplifies this gap. Although extensively studied in mammals, OXPHOS in other lineages has largely been inferred through sequence homology, an approach that may overlook lineage-specific components and propagate incorrect annotations. Here, we hypothesized that OXPHOS genes share characteristic transcriptional signatures that can be exploited for functional prediction. Using a curated set of 64 well-established OXPHOS genes, we combined supervised and unsupervised machine learning approaches to identify novel OXPHOS-associated genes in Caenorhabditis elegans . An ensemble of support vector machine, random forest, and k-nearest neighbors classifiers was trained on time-resolved bulk RNA-seq data using a novel informed bagging strategy and a two-round training scheme that incorporated genes annotated with limited evidence after an initial prediction round. In parallel, embryonic and adult single-cell RNA-seq datasets were used to infer co-expression networks and identify clusters enriched in known OXPHOS genes. Integrating both approaches yielded a high-confidence set of candidate genes supported by strong predictive performance on an independent test set. Several candidates lacked prior functional annotation. Functional validation of one top-ranked candidate, ril-1 , showed that a ril-1 mutant displayed significantly reduced oxygen consumption, consistent with a role for ril-1 in mitochondrial respiration.Our results demonstrate that integrating complementary machine learning strategies with transcriptomic data across multiple biological resolutions enables systematic discovery of genes associated with complex cellular processes.

Article activity feed