Likelihood-Based Inference and Model Selection for Stochastic Gene Expression in Probability-Generating-Function Space
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Selecting stochastic gene-expression models from single-cell counts requires accurate parameter inference and efficient model selection. Likelihood methods in count space can be costly when full stationary count distributions are unavailable, whereas approximate methods may lose accuracy. Probability generating functions (PGFs) offer a compact analytical alternative, but existing PGF workflows are generally not likelihood based and therefore rely on computationally intensive cross-validation. We develop a likelihood-based PGF framework for both tasks. Correlated empirical PGF values are used to construct a Gaussian quasi-likelihood for parameter inference and PGF-based Bayesian information criterion (BIC) for model selection. We show that the empirical PGF is exactly unbiased and that the parameter estimator is consistent, converges at the inverse-square-root sample-size rate, and is first-order asymptotically unbiased. For large samples and a uniquely preferred model, PGF-BIC selects the same model as leave-one-out cross-validation in PGF space.
Relevance to Life Sciences
Cell-to-cell variability in mRNA abundance contains information about the stochastic mechanisms regulating gene expression. Distinguishing constitutive production from transcriptional bursting and promoter switching is therefore important for interpreting single-cell RNA measurements, but conventional likelihood calculations and repeated cross-validation can be computationally expensive when many genes and candidate models are considered. The proposed PGF-BIC framework provides a scalable approach for parameter inference and model selection directly from single-cell count data. Its application to MERFISH nuclear and cytoplasmic mRNA counts enables gene-wise comparison of delayed Poisson and delayed Telegraph models and produces classifications that can be compared with those obtained by tenfold PGF cross-validation. The method supports efficient screening of stochastic gene-expression models while recognizing that selection of a statistical model does not by itself establish a unique molecular mechanism.
Mathematical Content
The empirical probability generating function (PGF), evaluated at a fixed set of collocation points, is represented as one correlated m -dimensional summary vector. A multivariate central limit theorem motivates a covariance-aware Gaussian quasi-likelihood whose maximizer is the weighted minimum-distance estimator used throughout the paper. Exact unbiasedness of the empirical PGF is proved, together with consistency, root- n asymptotic normality, and first-order asymptotic unbiasedness of the parameter estimator under the stated identification, smoothness, and uniform-integrability conditions. A Laplace approximation to the integrated quasi-likelihood yields PGF-BIC, comprising the fitted quadratic discrepancy and the penalty p log n . For a finite candidate set, regularity and positive separation in both count and fixed-grid PGF spaces ensure that PGF-BIC and count-space BIC asymptotically select the same smallest correct candidate. Leave-one-out (LOO) cross-validation is also formulated entirely in PGF space. With a uniquely separated population minimizer, PGF-LOO and PGF-BIC select the same candidate with probability tending to one.