HiDAC: A Hierarchical Dictionary-Aided Compression Framework for Genomic Sequences

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

The rapid increase in genomic data generated by next-generation sequencing (NGS) has created a strong demand for efficient DNA compression methods that can reduce storage requirements and support downstream sequence analysis. Existing approaches such as statistical, reference-based, and dictionary-based techniques provide good compression performance, but most of them are primarily designed for storage and require full decompression before any analytical operation can be carried out. In this work, a hierarchical dictionary-based DNA compression framework, HiDAC (Hierarchical Dictionary-Aided Compression) is proposed that can perform efficient compression and compresseddomain analysis of genomic sequences. The method identifies recurring patterns of variable length through an iterative cost-guided substitution strategy and represents them using reusable tokens. For efficient encoding, the generated dictionary is integrated with an Aho–Corasick based multipattern matching scheme followed by context-modelled arithmetic coding of the tokenised sequence. The proposed framework was evaluated on multiple genomic datasets, including Homo sapiens, Escherichia coli, and Mycobacterium tuberculosis, and was compared with general-purpose compressors as well as specialized genomic compression methods such as NAF and HMG. Experimental results showed that the proposed method achieved superior or competitive compression performance compared with the other methods across the tested datasets. In addition, substring-frequency analysis performed directly on the tokenised representation was on average 2.6–2.8× faster than the same analysis performed on the original sequence representation, across query lengths ranging from 13 to 466 bases and across two genomes of differing size and repeat structure. These results indicate that the proposed framework is effective not only for DNA compression but also for enabling computation directly on semi-compressed (tokenised) genomic data.

Article activity feed