A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 102 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
Rationale
Biological sequences exhibit a hierarchical organization that parallels the structure of natural language. At the residue level, amino acids function as functional tokens : catalytic residues behave like verbs that drive biochemical actions, hydrophobic residues form noun-like structural cores , and regulatory residues act as modifiers that tune activity. Short sequence motifs correspond to phrase-like units , domains operate as clause-level structures , and full proteins form coherent sentences within the broader paragraphs of cellular pathways. This linguistic analogy provides a natural justification for applying semantic and syntactic modeling frameworks to protein sequences.
By treating sequence elements as tokens embedded within a grammar-like system , we can analyze variant effects, domain interactions, and regulatory motifs using tools originally developed for natural language understanding . Conceptualizing proteins as linguistic constructs enables the use of semantic embeddings , contextual encoders , and hierarchical attention mechanisms to capture dependencies among residues, motifs, and domains that are not apparent from primary sequence alone. This perspective also offers a principled framework for interpreting model behavior: attention to serine or tyrosine residues corresponds to verb detection , clustering of hydrophobic residues reflects noun-like structural cores , and interactions among domains map to clause-level syntax .
In natural language, meaning arises not only from individual words but from their grammatical relationships , contextual dependencies , and hierarchical composition . Proteins exhibit analogous properties. Catalytic residues act as verbs that initiate biochemical reactions; hydrophobic cores provide the structural nouns; regulatory residues function as modifiers; and flexible linkers serve as conjunctions that connect functional units. Domains behave as clauses whose interactions determine the overall syntactic behavior of the protein. Mutations disrupt this grammar in predictable ways—ranging from minor spelling changes (missense variants) to truncated sentences (nonsense mutations) and syntactic collapse (frameshifts). This linguistic framework therefore supports the development of interpretable, biologically grounded machine-learning models capable of capturing the multi-level dependencies that govern protein function and variant impact.