Hō‘ike: A Joint-Embedding Predictive Architecture for Transcriptome Data Generation with Diffusion Models

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

In biomarker discovery, access to sufficient quantities of condition-specific transcriptomic data is often limited by cohort size, privacy concerns, and domain shift between normal and condition populations. Generative modeling can augment scarce cohorts and probe distributional transitions. Furthermore, synthetic transcriptome generation can support differential expression analyses, machine learning, privacy-preserving data sharing, benchmarking, and hypothesis generation in translational bioinformatics workloads in fields such as oncology.

Here, we present Hoike , a framework that combines a crossdomain Joint-Embedding Predictive Architecture (JEPA) with a latent diffusion model to generate condition-specific bulk transcriptomes from a normal reference context. In Hoike , normal tissue profiles provide continuous conditioning signals, while the model learns disease-linked shifts in latent space and reconstructs gene-level expression in log 2 (TPM+1) space. The implementation supports paired normal-condition training, tissuealigned conditioning, and constrained non-negative decoding for biologically valid outputs. We describe the architecture, objective design, and evaluation protocol used in this work across GTEx-derived normal references and multiple TCGA condition cohorts as a case study. This serves as the technical specification of the Hoike framework and its reproducible analysis workflow.

Article activity feed