Information Entropy of Biological Data: An Empirical Analysis
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Information entropy and its relatives—mutual information, relative entropy, and the maximum-entropy principle—are used pervasively across biology, yet they are notoriously difficult to estimate from finite samples, and the size of the resulting bias is rarely reported alongside the biological conclusion it supports. We present a unified empirical study that measures these quantities across six biological data modalities— genomic DNA, protein multiple-sequence alignments, transcription-factor binding sites, bulk-expression gene networks, single-cell transcriptomes, and microbiome communities—using a common suite of six estimators (plug-in, Miller–Madow, Chao–Shen, James–Stein shrinkage, NSB, and 𝑘-nearest-neighbour), and quantifies how much of each measured value is biology and how much is a sampling artifact. On synthetic controls with known ground truth the naive plug-in estimator is badly biased—underestimating a 64-symbol entropy by 0.66 bits at 𝑁=64 samples and reporting 0.56 bits of spurious mutual information between independent variables—while bias-corrected and 𝑘-NN estimators recover the truth to within a few hundredths of a bit. Carrying the corrected estimators to real data, we reproduce and quantify the canonical results of the field: DNA entropy rates of 1.87–1.91 bits/base with a genuine long-range redundancy that only repeat-aware compression exposes; the period-three mutual-information signature of coding sequence; Schneider’s 𝑅sequence ≈ 𝑅frequency law across 62 motifs (𝑟=0.98); a monotone improvement in protein contact prediction from raw mutual information (precision@𝐿 = 0.25) through average-product correction (0.42) to pseudolikelihood direct-coupling analysis (0.67); a 2–4 point gain in gene-network inference from replacing plug-in mutual information with shrinkage or 𝑘-NN estimators; a transcriptomic-entropy ordering of differentiation potency (Spearman 𝜌= − 0.87); and a saturating single-cell foundation-model loss that converges to an irreducible per-gene entropy floor of ≈2.3 bits. We assemble the corrected values into a cross-domain “entropy atlas” and show that the sign and magnitude of the finite-sample bias are as predictable as they are consequential: entropies are systematically underestimated, mutual information and information content overestimated, and in the most undersampled regimes the bias can exceed the biological signal. We argue that reporting entropy without its estimator and sample size is no longer defensible, and that modern self-supervised models are best read as large conditional-entropy estimators subject to the same century-old sampling constraints.