Rethinking respiratory disease forecasting: temporal heterogeneity between surveillance predictors and outcomes drives forecast instability
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Since the COVID-19 pandemic, forecasting hubs and non-traditional respiratory disease surveillance streams have become increasingly common. However, many forecasting approaches assume that relationships between surveillance predictors and disease outcomes remain stable over time and that incorporating additional historical data will improve forecast performance. To evaluate these assumptions in a real-world setting, we developed and evaluated forecasts of SARS-CoV-2 and influenza hospitalizations in Utah using syndromic surveillance, test positivity, and wastewater data. Rather than identifying a single, best-performing model, we examined whether relationships between surveillance predictors and hospitalization outcomes remained stable across seasons and whether longer historical training periods consistently improved forecast accuracy. Relationships between surveillance predictors and hospitalizations varied substantially by pathogen and season. Analyses using pooled data across multiple years suggested strong positive correlations between predictors and outcomes, but these aggregated patterns often obscured weak or negative correlations observed during SARS-CoV-2 variant waves and influenza seasons. Forecast performance similarly varied over time. Models that performed well during some seasons, transmission phases, or under certain training strategies frequently performed worse than benchmark models in others. Training on additional historical data generally reduced forecast accuracy, though this varied by disease and transmission phase. Forecasting groups should prioritize continual evaluation of surveillance predictors, adaptive strategies, and diverse ensembles, rather than relying on a single model, data stream, or historical training framework each year.
AUTHOR SUMMARY
Respiratory disease forecasting hubs and novel data streams have become integral parts of infectious disease surveillance and public health decision-making since the COVID-19 pandemic. Many forecasting groups assume that adding more historical data will improve model performance and that relationships between surveillance predictors, such as emergency department visits or wastewater, and hospitalizations will remain stable over time. We evaluated these assumptions using forecasts of SARS-CoV-2 and influenza hospitalizations in Utah. We found that relationships between surveillance predictors and hospitalizations varied across SARS-CoV-2 variants, influenza seasons, and periods of increasing and decreasing transmission. Forecast performance also varied considerably, with models that performed well in some seasons often performing poorly in others. Public health groups should continually evaluate the utility of surveillance predictors in real-time and prioritize adaptable, diverse modeling approaches.