How can spectrum bias impact expectations of test performance: a secondary modeling study estimating novel swab-based test outcomes across populations
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
New swab-based near-point-of-care (NPOC) tests offer a potentially lower-cost, simpler alternative to Xpert MTB/RIF Ultra (Xpert-Ultra) testing for tuberculosis (TB) diagnosis. However, it is critical to understand how their performance may differ across diverse populations to inform programmatic rollout and real-world clinical decision-making.
Methods
We modeled the performance of testing sputum swabs with the MiniDock MTB assay and the Xpert MTB/RIF (Xpert) using positive percentage agreement (PPA) with Xpert-Ultra, disaggregated by semi-quantitative grade. PPA for MiniDock MTB was derived from a random-effects meta-analysis of three diagnostic accuracy studies. PPA for Xpert was derived from data from an early diagnostic study. PPA estimates were applied to 1,248 positive Xpert-Ultra test results from Phase 1 of the Start4All study across seven countries, disaggregated by facility-based (n=1,033) and community-based (n=215) participant recruitment, comparing a single overall PPA to an Xpert Ultra semi-quantitative grade-stratified model (from Trace to High).
Results
Pooled overall PPA was 84.8% (95% CI: 67.1–93.8%) for MiniDock MTB and 93.1% (90.3–95.2%) for Xpert. Grade-stratified modeling revealed lower PPAs at ‘Very Low’ and ‘Trace’ semi-quantitative grades: MiniDock MTB 56.5% and 33.8%; Xpert 66.7% and 23.1%, respectively. When grade-stratified estimates were applied to the Start4All data, both tests performed similarly, missing 241/1,248 and 228/1,248 positive Xpert-Ultra results, respectively. The single-value model overestimated performance most markedly in community settings, predicting 10.8% and 19.1% more positive results than the grade-stratified model for MiniDock MTB and Xpert, respectively.
Conclusion
MiniDock MTB sputum swabs perform comparably to Xpert when assessed using grade-stratified modeling. Both tests are likely to miss a greater proportion of people with Xpert-Ultra positive results in community settings, where paucibacillary disease is more common. Using single point estimates for diagnostic accuracy can substantially overestimate real-world performance, highlighting the importance of evaluating diagnostics across the full spectrum of TB disease.