A Multi-Layered Plagiarism and AI-Generated Content Detection Framework Integrating BERT-BiLSTM-Attention Encoding with Stylometric Analysis
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
The proliferation of large language models (LLMs) and the persistent challenge of academic dishonesty necessitate detection systems capable of identifying not only verbatim copying but also paraphrased, mosaic, and machine-generated text. Existing plagiarism detection tools predominantly rely on lexical overlap metrics, which prove insufficient against sophisticated paraphrasing and AI-authored content. This paper presents a comprehensive detection framework that integrates a Bidirectional Encoder Representations from Transformers–Bidirectional Long Short-Term Memory–Multi-Head Attention (BERT-BiLSTM-Attention) architecture with Sentence-BERT (SBERT) semantic embeddings, n-gram fingerprinting, and a twelve-feature stylometric AI-writing detector. The system performs four distinct detection tasks: exact-match identification via sequence alignment, paraphrase detection through cosine similarity over contextual embeddings, mosaic or patchwork plagiarism recognition using hashed n-gram overlap, and AI-generated content flagging based on statistical heuristics including type-token ratio, burstiness, transition-word density, and adjacent-sentence similarity. Additionally, Density-Based Spatial Clustering of Applications with Noise (DBSCAN) clustering over sentence-level embeddings enables automated writing-style segmentation, facilitating the identification of inconsistent authorship within a single document. A rigorous validation suite employing synthetic corpora with known ground truth evaluates the pipeline across unit, integration, edge-case, and consistency dimensions. Experimental results demonstrate stable detection performance across diverse document types, confirming the framework's robustness and its potential as a scalable, multi-faceted integrity assessment tool for academic and editorial contexts.