Interpreting the CTCF-mediated sequence grammar of genome folding with AkitaV2

Paulina N. Smaruj
Fahad Kamulegeya
David R. Kelley
Geoffrey Fudenberg

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Interphase mammalian genomes are folded in 3D with complex locus-specific patterns that impact gene regulation. CTCF (CCCTC-binding factor) is a key architectural protein that binds specific DNA sites, halts cohesin-mediated loop extrusion, and enables long-range chromatin interactions. There are hundreds of thousands of annotated CTCF-binding sites in mammalian genomes; disruptions of some result in distinct phenotypes, while others have no visible effect. Despite their importance, the determinants of which CTCF sites are necessary for genome folding and gene regulation remain unclear. Here, we update and utilize Akita, a convolutional neural network model, to extract the sequence preferences and grammar of CTCF contributing to genome folding. Our analyses of individual CTCF sites reveal four predictions: (i) only a small fraction of genomic sites are impactful; (ii) impact is highly dependent on sequences flanking the core CTCF binding motif; (iii) core and flanking nucleotides contribute largely additively to the overall impact of a site; (iv) sites created as combinations of different core and flanking sequences have impacts proportional to the product of their average impacts, i.e. they are broadly compatible. Our analysis of collections of CTCF sites make two predictions for multi-motif grammar: (i) insulation strength depends on the number of CTCF sites within a cluster, and (ii) pattern formation is governed by the orientation and spacing of these sites, rather than any inherent specialization of the CTCF motifs themselves. In sum, we present a framework for using neural network models to probe the sequences instructing genome folding and provide a number of predictions to guide future experimental inquiries.

Version published to 10.1371/journal.pcbi.1012824
Feb 4, 2025
Version published to 10.1101/2024.08.01.606065 on bioRxiv
Aug 4, 2024

Cross–Cell-Line Conservation-Resolved Interactions Reveal a Stepwise Cascade of CTCF Loop Formation

This article has 3 authors:
1. Maryam Mirabolghasemi¹
2. Mohammad Hossein Karimi-Jafari¹
3. Ali Mohammad Banaei-Moghaddam²
This article has no evaluationsLatest version Feb 24, 2026
A proteocistromic atlas of 216 human disease-relevant transcription factors

This article has 17 authors:
1. Gong-Hong Wei
2. Zixian Wang
3. Zenglai Tan
4. Xiaonan Liu
5. Guowen Duan
6. Peng Zhang
7. Wenjie Xu
8. Binjie Luo
9. Longguang Qin
10. Yuehong Yang
11. Shuangshuang Ma
12. Xiayun Yang
13. Matias Kinnunen
14. Iftekhar Chowdhury
15. Qin Zhang
16. Aki Manninen
17. Markku Varjosalo
This article has no evaluationsLatest version Mar 12, 2026
High-resolution binding data of TFIID and cofactors show promoter-specific differences in vivo

This article has 4 authors:
1. Julia Zeitlinger
2. Sergio García-Moreno Alcántara
3. Simon Bourdareau
4. Melanie Weilert
This article has no evaluationsLatest version Jan 30, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Cross–Cell-Line Conservation-Resolved Interactions Reveal a Stepwise Cascade of CTCF Loop Formation

A proteocistromic atlas of 216 human disease-relevant transcription factors

High-resolution binding data of TFIID and cofactors show promoter-specific differences in vivo