Cells are not replicates: permutation calibration of the unit-of-analysis error in a pooled sleep single-cell design (GSE137665)

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background. In single-cell RNA sequencing the number of cells is not the number of independent observations for a between-animal intervention. When animals are pooled before capture, the animal label is destroyed and the between-animal variance component becomes unestimable. Methods. We re-analysed GSE137665 (mouse brainstem, cortex and hypothalamus across normal sleep, 12-h sleep deprivation and recovery sleep). The deposited matrix holds 29,571 cells × 18,634 genes in nine libraries, each a suspension pooled from three animals. We fitted the same region-adjusted model for the sleep-deprived vs normal-sleep contrast at two units of analysis - individual cells (df = 29,566) and library means (n = 9, df = 4) - and calibrated both by permuting condition labels within brain region (216 assignments). We also decomposed the variance into within- and between-library components, and audited the study's RNAscope validation layer and its own deposited differential-expression output. Results. Under the permutation null the cell-level test returned a mean of 6,324 genes at p < 0.05 (median 6,489), an empirical type-I error rate of 37.7% at a nominal α = 0.05; the library-level test returned 3.0%. The observed cell-level result (4,688 genes) lies below its own null mean (permutation p = 0.759). The Kruskal-Wallis test used by the source study gives a median per-gene p-value of 1.9 × 10⁻⁴ , against 0.5 for a calibrated test. At the library level no gene passes FDR < 0.1 or Bonferroni; the smallest effect detectable at 80% power (0.278) exceeds the 99th percentile of observed effects (0.113). Point estimates agree between the two units (r = 0.993), so the unit of analysis changes the inference, not the estimate. The RNAscope layer shows the same error: 10,419 − 21,227 counted cells per panel from three brains per group, with a critical between-brain ICC of 0.023–0.146. Conclusions. No valid population-level test can be constructed from these data, and because the defect lies in the design it cannot be repaired retrospectively. Sleep scRNA-seq reports should state n = animals, avoid pooling before capture, and calibrate cell-level tests by permutation whenever animals are few.

Article activity feed