TMO: ASYMMETRIC CROSS-MODAL ATTENTION FOR LEARNING CELL-STATE-DEPENDENT REGULATORY LAGS FROM SINGLE-CELL MULTIOMIC DATA

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Single-cell multi-omics technologies simultaneously measure chromatin accessibility (ATAC) and gene expression (RNA), providing a unique window into the temporal ordering of regulatory events during differentiation. However, most computational models treat the two modalities symmetrically, ignoring the directional relationship between chromatin and transcription, and existing lag-aware methods estimate a single global lag per gene, failing to capture cell-state-dependent dynamics.

Methods and Results

We introduce Temporal Multi-Omics (TMO), a deep learning framework that learns signed, cell-state-conditional regulatory lags (Δ τ ) using asymmetric cross-modal attention. TMO projects RNA and ATAC into 50 latent components each, tokenises each cell as a sequence of 100 tokens, and uses a two-pass transformer in which a data-driven lag prior – derived from a sliding-window cross-correlation function – directly biases attention asymmetrically. On four independent 10x Multiome datasets (mouse brain, human brain, mouse kidney, human PBMC), the asymmetric model achieves Lag Concordance Scores (LCS) of 0.988–0.999, compared to 0.048–0.108 for an architecturally identical symmetric baseline. A stratified 80/20 held-out experiment confirms that the learned component-lag ordering generalises to unseen cells (held-out LCS 0.85–0.99). Clustered Δ τ heatmaps show positive Δ τ (ATAC-led priming) in early pseudotime and negative Δ τ (RNA-led, activity-dependent regulation) in late pseudotime; the ATAC-RNA correlation heatmap exhibits a U-shaped pattern indicative of developmental decoupling. Components with the most positive Δ τ are enriched for chromatin organization and stem cell differentiation (FDR < 0.05), while those with the most negative Δ τ are enriched for synaptic signalling and immune activation. Ablating the cell-state information from the lag predictor reduces the LCS and collapses per-component temporal dynamics (KS p ≤ 0.039 in all four tissues), proving that TMO’s dynamic lag patterns depend on cell-state conditioning. Independent ChIP-seq validation for four transcription factors (PAX5, Pax6, ASCL1, Hnf4 α ) confirms highly significant separation between target genes and expression-matched background ( p < 10 −4 in all cases). Two Multiome Perturb-seq screens provide causal validation: SMARCB1 knockout shows a directional trend (1.5-fold target shift, p = 0.056, n = 147 perturbed cells), and SMARCE1 knockout reaches statistical significance ( p = 0.0089, n = 3,394 perturbed cells). Gene-level cross-correlation independently validates that the regulatory lag signal is present in the raw data, and TMO further identifies rare, statistically significant biphasic gene programs where the regulatory direction reverses across pseudotime.

Conclusions

T MO is the first method to make regulatory lag a learnable, cell-state-conditional, and architecturally encoded parameter. It is scalable, interpretable, and open-source, providing a powerful tool for studying regulatory timing in development, disease, and perturbation screens.

Highlights

  • TMO learns signed, cell-state-conditional regulatory lags (Δ τ ) from paired ATAC+RNA data.

  • Asymmetric cross-modal attention biases the model towards the direction of regulation (ATAC leads RNA or RNA leads ATAC).

  • The asymmetric model achieves Lag Concordance Scores > 0.98, far outperforming a symmetric baseline (< 0.11).

  • Stratified 80/20 held-out experiments across all four tissues show that the learned component-lag ordering transfers to unseen cells (held-out LCS 0.85–0.99).

  • Perturb-seq screening detects perturbation-induced shifts in regulatory timing: SMARCB1 knockout shows a directional trend ( p = 0.056), and SMARCE1 knockout reaches significance ( p = 0.0089), demonstrating causal relevance.

  • Independent ChIP-seq validation across four transcription factors and tissues confirms that TMO-derived lags distinguish physically bound genes from expression-matched background ( p < 10 −4 in all cases).

  • Removing the cell embedding from the LagMLP causes a substantial drop in LCS and collapses the per-component temporal dynamics, proving that TMO’s cell-state conditioning is essential.

  • TMO is open-source, scalable, and ready for application to any 10x Multiome dataset.

  • Gene-level cross-correlation confirms the lag signal is intrinsic to the data and is 3.7–4.7× more variable than TMO’s denoised component profiles.

  • TMO identifies statistically significant biphasic gene programs whose regulatory direction reverses across pseudotime, a new dynamic inaccessible to existing scalar methods.

  • 1.

    Figure 1: Graphical Abstract

    Figure 1:

    Graphical Abstract | TMO – Temporal Multi-Omics.

    TMO uses asymmetric cross-modal attention in a two-pass transformer to learn signed, cell-state-conditional regulatory lags (Δ τ ) between chromatin accessibility (ATAC) and gene expression (RNA) from paired single-cell multi-omic data. A data-driven lag prior derived from a sliding-window cross-correlation function biases ATAC-to-RNA attention, enabling the model to capture ATAC-led priming (positive Δ τ , red) early in differentiation and RNA-led, activity-dependent regulation (negative Δ τ , blue) late in differentiation. Across four tissues, the asymmetric model achieves Lag Concordance Scores (LCS) of 0.988–0.999, while a symmetric baseline collapses to <0.11. Independent ChIP-seq and Perturb-seq validations confirm that the learned lags reflect genuine transcription factor binding and respond to genetic perturbations. TMO is open-source, scalable, and available as a Scanpy-like Python package (tmopy). Made with BioRender.

    Article activity feed