Can Machine Learning Identify Meaningful Patterns in DNA Sequences? A Computational Study of DNA Motifs Associated with Transcription-Factor Binding

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

This independent computational study examines whether machine-learning methods can distinguish CTCF-associated DNA sequences from a controlled background using short nucleotide patterns (k-mers). The analysis draws on a publicly available UniBind dataset for human CTCF in K562 cells, linked to ENCODE experiment ENCSR000EGM and JASPAR motif profile MA0139.1. The dataset comprises 46,846 PWM-derived sequences of 19 nucleotides each. Negative examples were generated through dinucleotide shuffling of the positive sequences, preserving local base-pair composition while removing higher-order sequence structure. Each sequence was represented using normalized 3-mer and 4-mer frequencies, yielding 320 numerical features. Logistic Regression and Random Forest classifiers were trained on an 80/20 held-out split and evaluated using accuracy, precision, recall, F1-score, ROC-AUC, and confusion-matrix analysis. Random Forest achieved the stronger performance (92.53% accuracy; ROC-AUC 0.979) compared with Logistic Regression (86.49% accuracy; ROC-AUC 0.938). These results indicate that short sequence-composition features carry substantial discriminative information relative to a dinucleotide-shuffled background. Because the negative class was derived from the positive sequences themselves, the findings should not be read as a general predictor of CTCF binding across arbitrary genomic DNA. Rather, the project's principal contribution is a transparent, reproducible workflow linking molecular biology, sequence analysis, and applied machine learning — the kind of end-to-end pipeline typically introduced at the undergraduate level.

Article activity feed