How Predictable Was the 2026 Bundibugyo Virus Disease Outbreak? A Rolling-Origin Evaluation of Short-Term Forecast Models, an Empirically Recalibrated Bayesian Model, and a Data-Driven Baseline–Bayesian Ensemble, Using Daily Surveillance Data

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

The 2026 outbreak of Bundibugyo virus disease in the Democratic Republic of the Congo generated the largest recorded epidemic caused by Bundibugyo virus and displayed unusually rapid early growth. Early scenario projections warned that the outbreak could become one of the largest Ebola-family epidemics on record, but were calibrated to assumed mortality totals and intervention scenarios rather than evaluated repeatedly against subsequently observed daily surveillance data. We assessed how accurately the outbreak could have been forecast in real time using routinely published national situation reports, whether a simple ensemble of two complementary models improved on either alone, and how the resulting forecasts and their uncertainty should be interpreted for operational capacity planning.

Methods

We conducted a retrospective, pseudo-prospective rolling-origin forecasting study using daily cumulative confirmed cases and deaths reported in national situation reports from 14 May to 20 July 2026 (68 calendar days; 61 numeric national reports). Every date with a newly reported national total became an eligible forecast origin once at least seven numeric observations were available (55 origins). At each origin, all later observations were withheld and two prespecified operational models were refitted using only data available at that time: a recent seven-increment baseline and a Bayesian negative-binomial surveillance-maturity model fitted by full Markov Chain Monte Carlo on complete available history. A Gompertz growth model was also fitted identically at every origin as a phenomenological benchmark. Forecasts were generated for 7-, 14-, and 21-day horizons and evaluated only when an observation existed on the exact target date. The Bayesian model’s forecasts were additionally corrected by a prospective, quality-gated empirical recalibration using only errors already realised at earlier, forecast-maturity-qualified origins. We further evaluated a data-driven ensemble that pooled the baseline’s bootstrap trajectories with the recalibrated Bayesian model’s posterior predictive trajectories, with horizon- and outcome-specific mixing weights selected by a strictly prospective, expanding-origin procedure using an explicit asymmetric operational loss function.

Findings

A statistical and externally corroborated surveillance-maturity discontinuity was identified on 28 May 2026 (robust z-score 13·0 against the trailing week). The Gompertz model had the poorest 80% predictive-interval coverage of all candidates at every horizon and outcome (25·9–34·1% for cases, 0·0–23·5% for deaths) and is reported as tested and excluded. Recalibration corrected a directionally consistent Bayesian underprediction tendency and is the primary reported Bayesian result; the correction was applied at 24 of 55 origins for the 21-day horizon (never at 28 days in earlier work), because 21-day forecasts reach the forecast-maturity threshold (mature history ≥ forecast horizon) from the 18 June origin onward, a genuinely more useful operational horizon than 28 days. Recalibration improved case coverage substantially (7-day 80% coverage 60·0%→80·0%; 14-day 41·0%→76·9%) and improved death forecasts on both accuracy and calibration (7-day median absolute percentage error 10·3%→7·7%; 80% coverage 33·3%→95·6%). The baseline–Bayesian ensemble’s evidence-selected weighting differed materially by outcome: for cases, the operationally and distributionally optimal blend was baseline-heavy (w≈0·8–0·9 toward the baseline) at every horizon; for deaths, the optimal blend was close to parity (w≈0·5–0·6), and the ensemble improved mean weighted interval score over both individual component models at every horizon (e.g., 7-day: baseline 24·8, Bayesian 20·5, ensemble 17·9).

Interpretation

Routine SitRep data supported useful short-term forecasting, but no single model was best on every criterion, and a data-driven combination of two complementary, individually weaker-in-some-respect models measurably outperformed either alone for one of two outcomes evaluated. Three distinct kinds of maturity determine whether a forecast at a given horizon can be trusted: epidemic maturity, surveillance maturity, and forecast maturity; the last of these is the strongest and most model-agnostic finding in this paper, and motivates reporting 21- rather than 28-day forecasts as the longest routine operational horizon. We report an explicit, reproducible, and adjustable framework — not a fixed recommendation — for combining simple and complex models and for communicating the resulting uncertainty to non-specialist operational decision-makers.

Funding: none received specifically for this analysis.

Author Summary

When an outbreak is growing, planners need to know not just how many cases to expect, but how much to trust that number. We tested, in a fully retrospective and fair way, how well three forecasting methods — a simple recent-trend method, a saturating growth curve, and a more sophisticated probabilistic model — could have predicted the 2026 Bundibugyo virus disease outbreak in the Democratic Republic of the Congo one, two, and three weeks ahead, using only the information that would genuinely have been available at each point in time. The saturating growth-curve model performed worst on every measure and is not recommended. The simple trend method gave the most accurate single-number forecasts for confirmed cases, but its stated uncertainty ranges were frequently wrong. The probabilistic model’s ranges were much better calibrated once we corrected a systematic bias using only past, already-known forecast errors — a technique we describe in enough detail that others can reuse it on their own outbreaks. Combining the simple and probabilistic methods, in carefully validated proportions that differ for cases and for deaths, produced better forecasts than either method alone for deaths specifically. We also explain, in plain language, what a forecast’s uncertainty range does and does not mean for a non-statistician planning hospital beds, staff, or supplies, and we show that three-week forecasts are far more trustworthy than four-week ones simply because the outbreak had not yet generated enough mature surveillance history to test a four-week forecast fairly — a checkable data limitation, not a permanent weakness of any model.

Article activity feed