Efficient Grammar Compression via RLZ-based RePair

Rahul Varki
Travis Gagie
Christina Boucher

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Among grammar-based compression techniques, RePair is a notable offline encoding scheme known for its simplicity and powerful combinatorial properties, producing compact grammars by repeatedly replacing the most frequent adjacent pairs of symbols, known as bigrams. However, RePair's memory usage scales poorly with input size, as it loads the entire text into memory. In contrast, Relative Lempel-Ziv (RLZ) parsing offers a scalable and lightweight online encoding scheme that losslessly represents a text in terms of phrases that refer to a reference string, but it often fails to expose deeper structural patterns. We introduce an algorithm that produces a RePair grammar from the RLZ parse of the input, leveraging the strengths of both methods. Our method, RLZ-RePair, performs bigram replacements systematically, preserving the integrity of the RLZ phrases throughout the RePair iterations. When the reference is well chosen, our method achieves the same grammar as standard RePair while significantly reducing both memory usage and the number of bigram replacements. In particular, we show that RLZ-RePair uses significantly less memory than RePair, which required between 18% and 480% more memory across different data sets. To our knowledge, RLZ-RePair is one of the first scalable methods that constructs exact RePair grammars, resulting in a grammar-based compressor that is both practical for large datasets and faithful to the theoretical elegance of RePair.

Version published to 10.1101/2025.07.22.666196 on bioRxiv
Jul 26, 2025

Text Distance from Nested and Hierarchical Repetitions: A Compression-Based Perspective

This article has 7 authors:
1. Xiaojun Hu
2. Jing Wang
3. Jingwen Zhang
4. Fengyao Zhai
5. Xiao Xie
6. Zengru Di
7. Yu Liu
This article has no evaluationsLatest version Jun 12, 2025
Ladderpath: An Efficient Algorithm for Revealing Nested Hierarchy in Sequences

This article has 9 authors:
1. Jingwen Zhang
2. Xiao Xie
3. Xiaodong Deng
4. Jing Wang
5. Xiaojun Hu
6. Yiping Wang
7. Hu Zhu
8. Fengyao Zhai
9. Yu Liu
This article has no evaluationsLatest version Jul 21, 2025
Toward Efficient and Faithful Reasoning in Large Language Models

This article has 3 authors:
1. Lukas Schneider
2. Anna Muller
3. Mareike Gerhardt
This article has no evaluationsLatest version Jul 18, 2025

Listed in

Abstract

Article activity feed

Related articles

Text Distance from Nested and Hierarchical Repetitions: A Compression-Based Perspective

Ladderpath: An Efficient Algorithm for Revealing Nested Hierarchy in Sequences

Toward Efficient and Faithful Reasoning in Large Language Models