The Hidden Prior: Variance Constraints Under Data Augmentation
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Data augmentation is now a standard device across capture–recapture and occupancy analysis: adding a fixed number M of all-zero encounter histories replaces a model of unknown dimension with one of fixed dimension. Although M is often treated as a computational tuning choice, it also specifies a finite superpopulation and hence a binomial support constraint on the number of undetected individuals. In a Bayesian implementation that constraint appears as an induced prior; in a likelihood implementation it is the same finite-support assumption reached by another route. That the Bernoulli specification for the inclusion indicators induces a binomial prior on abundance is established (Schofield & Barker 2014); our concern is what that choice costs in estimated uncertainty. We develop the argument using a simple closed-population abundance estimation problem. We show that augmented occupancy and Huggins conditional-likelihood analyses give numerically identical point estimates of once M is sufficiently large. Their uncertainty estimates, however, need not agree. We distinguish two sources of discrepancy. First, when M is small relative to the number of undetected individuals, the finite binomial ceiling truncates the likelihood or posterior and suppresses uncertainty. Second, once that ceiling no longer binds, Taylor-series (Delta-method) approximations still understate variance, because the quantity of interest is a strongly non-linear function of the estimated parameters and local linearization does not reproduce its curvature. Gauss–Hermite quadrature on the unconstrained logit scale recovers much of the shortfall and approaches the MCMC posterior benchmark, though a small residual remains that does not close as M grows, reflecting the distinction between asymptotic likelihood theory and finite-sample Bayesian inference. Neither mechanism is peculiar to abundance estimation: the first follows from the augmented representation itself, the second from any derived quantity that is a non-linear function of estimated parameters. We close with framework-specific guidance for choosing M.