Digital Acoustic Phenotypes of Overnight Respiratory Activity
This article has been Reviewed by the following groups
Listed in
- Evaluated articles (PREreview)
Abstract
Purpose: Respiratory-event rates quantify event frequency but do not describe how respiratory-related acoustic activity is expressed across an overnight recording. We investigated whether a score-defined digital acoustic representation captures information not reducible to annotated respiratory-event rate.Methods: We analyzed 32 overnight home respiratory polygraphy recordings from the APSAA dataset, including audio, respiratory annotations, and oxygen saturation. Twenty-eight recordings yielded at least one qualifying acoustic episode (8,512 episodes). Five acoustic features—episode density, relative intensity, duration, long-event occurrence, and temporal distribution—formed a composite score defining three acoustic phenotypes using fixed heuristic thresholds. Temporal compaction quantified short inter-episode intervals and was adjusted for episode quantity.Results: Scores were computationally reproduced with identical phenotype assignment in 28/28 recordings (mean absolute difference 0.0011; maximum 0.0046). Two recordings with comparable annotation-derived respiratory-event rates (30.45 vs 33.38 events/h; 217 vs 232 annotated events) and mean SpO₂ (91.75% vs 91.61%) contained 382 vs 5 acoustic episodes and scores of 0.685 vs 0.145. Load-adjusted temporal compaction differed among phenotypes (H=8.192, p=0.0166, ε²=0.248). The effect persisted after removing the historical 120-s interval exclusion (H=8.108, p=0.0174, ε²=0.244); residual association with episode count remained. The low- versus intermediate-score contrast was robust, whereas the high-score contrast was less stable.Conclusions: Overnight respiratory audio can provide reproducible digital acoustic phenotypes describing the amount and temporal organization of detected acoustic activity beyond annotated event frequency. These are signal representations, not clinical phenotypes or sleep stages. Repeated-night studies are required to establish within-person stability, transitions, and clinical meaning.
Article activity feed
-
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22789060.
Summary of main findings
This exploratory study reanalyzes 32 overnight home recordings (audio + respiratory polygraphy) from the APSAA dataset for obstructive sleep apnea. The premise: a respiratory-event rate (AHI-like) summarizes an entire night as one number but discards how respiratory-related acoustic activity is distributed over time. The author built a "digital acoustic phenotype" from five acoustic features (episode density, relative intensity, duration, late long-event occurrence, temporal distribution) combined into a composite score classifying each recording as High/Intermediate/Low.
Core findings: (1) The score was computationally reproduced with near-perfect fidelity in …
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22789060.
Summary of main findings
This exploratory study reanalyzes 32 overnight home recordings (audio + respiratory polygraphy) from the APSAA dataset for obstructive sleep apnea. The premise: a respiratory-event rate (AHI-like) summarizes an entire night as one number but discards how respiratory-related acoustic activity is distributed over time. The author built a "digital acoustic phenotype" from five acoustic features (episode density, relative intensity, duration, late long-event occurrence, temporal distribution) combined into a composite score classifying each recording as High/Intermediate/Low.
Core findings: (1) The score was computationally reproduced with near-perfect fidelity in 28/28 recordings. (2) This acoustic representation does not simply mirror annotated respiratory-event rate — two recordings with nearly identical annotated rates (30.45 vs 33.38 events/h) contained 382 vs 5 acoustic episodes. (3) Load-adjusted "temporal compaction" differed across phenotypes (p=0.0166), but the effect was concentrated almost entirely in the Low-score group.
Field contribution: the core idea — that event frequency alone hides temporal organization — is methodologically reasonable and worth pursuing, but this is a very preliminary, small-sample exploratory analysis, not a demonstration of real biological phenotypes, as the author himself repeatedly and explicitly states throughout.
Major issues
Very small sample, especially the Low-score group: only 4 recordings drive the paper's headline statistical result (temporal compaction). Leave-one-out analyses show the Low-vs-High comparison is unstable (significant in only 1 of 4 omissions), while Low-vs-Intermediate is comparatively more robust.
Arbitrary, non-biological thresholds: the High/Intermediate boundary separation is just 0.004 around the 0.60 cutoff — smaller than the reconstruction error itself (max 0.0046). This means the two categories are essentially indistinguishable at their boundary and the threshold has no independent justification.
One component saturated in 75% of the sample: the most heavily weighted component (density, 30% weight) hit its ceiling in 21/28 recordings, meaning it contributed zero discriminating power across three-quarters of the data — undermining the composite score's claimed multidimensionality.
Detector not validated against ground truth: the acoustic segmentation rule was "inherited" from a prior exploratory analysis and never optimized or validated against the respiratory annotations. Four recordings yielded zero acoustic episodes despite substantial annotated respiratory event burden (one had 537 annotated events, second-highest in the dataset) — a fundamental detector failure mode that the paper acknowledges but doesn't resolve.
No clinical validation whatsoever: the "phenotypes" are explicitly stated to not represent sleep stages, OSA severity, or any clinical construct — this is appropriately cautious, but it also means the paper demonstrates a signal-processing pipeline is reproducible, not that it measures anything clinically meaningful.
Single night per participant: no within-person reproducibility, stability, or night-to-night variability can be assessed — a critical gap given documented substantial night-to-night variability in OSA itself (cited by the author).
Pending patent application disclosed at the end — worth noting as a potential conflict-of-interest consideration even though formally declared.
Minor issues
The Methods section is extremely dense with normalization formulas and rounding rules; a summary table of the scoring pipeline early on would aid readability.
Repeated caveating throughout ("this is not a clinical phenotype...") is appropriate but becomes somewhat repetitive — could be consolidated into one clear framing statement plus targeted reminders.
Figure 1B is visually informative but a supplementary table of per-component saturation rates would make the "component D is capped in 75%" point more quantitatively transparent.
The comparison between annotation-derived rate and acoustic score (Section 3.3) would benefit from a scatter plot matrix rather than prose-heavy pairwise anchor-recording narration.
AI-assistance disclosure is present and appropriately scoped (workflow/scripting/figures only, not analysis/interpretation) — good practice, no issue there.
Competing interests
The author declares that they have no competing interests.
Use of Artificial Intelligence (AI)
The author declares that they did not use generative AI to come up with new ideas for their review.
-