What Week 8 Knows: Forecasting Six-Month GLP-1 Outcomes

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Patients on GLP-1 medications lose very different amounts of weight, and most published prediction models include only patients who complete six months. That design omits everyone who disengages earlier, which is the majority of the cohort. We built a tool that includes patients who disengage and delivers useful predictions at the week-8 visit, where the clinical decision is actually made.

Methods

Beginning with 237,800 adults enrolled in a US telehealth GLP-1 program, we required a documented week-8 weight, a refill-confirmed dose, and reported ethnicity, yielding an analytic cohort of 22,538. We answered three questions: the patient’s likely six-month weight loss and our confidence in it; the probability of dropout before six months; and when weight loss plateaus. For the first, we fit a cubic in week-8 percent loss plus 16 covariates, with quantile-regression bands at the 10th and 90th percentiles for the prediction interval, checking fractional-logit and isotonic recalibration as alternatives. For the second, we fit a logistic regression and compared it to gradient boosting. For the third, we fit a per-patient exponential trajectory among patients with at least four weight observations. We trained on enrollments before 2024-07-01 and tested on later ones, compared completer outcomes to published RCTs, and tested the week-8 anchor (week 8 is the anchor itself) against measurements at weeks 2, 4, 6, 8, 10, 12, 16, and 20.

Results

Mean six-month weight loss in completers was 11.7% on semaglutide and 14.1% on tirzepatide, in line with STEP-1 and SURMOUNT-1. 1,2 Six-month disengagement was 66%. The prediction model reached test R 2 = 0.65 with a mean absolute error of 2.76 percentage points. Calibration was strong: calibration-in-the-large was 0.52 pp and the calibration slope was 0.96. The 80% quantile-regression interval covered 76% of test patients; the 95% interval covered 93%. The disengagement model reached test AUC 0.79, against 0.74 for gradient boosting. Median plateau time among engaged patients was 387 days, longer in lower-BMI tertiles. The week-8 anchor gave R 2 = 0.65, compared to 0.48 to 0.61 at earlier weeks and 0.67 to 0.91 at later weeks. We chose week 8 because 80% of slow responders reach their post-titration decision point at or before that visit. Two of twenty subgroup cells had reduced predictive accuracy; two more were too sparse to validate.

Conclusions

Observed week-8 weight loss is the strongest predictor of six-month outcome. The model’s accuracy (R 2 = 0.65, MAE 2.76 pp) is appropriate for calibrating expectations and identifying patients for the post-titration decision, but not precise enough to drive that decision on its own. Disengagement is predictable at week 8 with AUC 0.79. Engaged patients plateau at a median of 387 days. Week 8 is the earliest visit at which titration is mostly complete, accuracy is in a useful range, and the post-titration decision remains actionable; later anchors predict better but inform a decision that has already been made for most patients. The model is temporally (internally) validated but not yet externally validated, and because it was developed on a single platform it should be regarded as a recalibration target rather than a drop-in deployment elsewhere. The tool is published as a public web calculator to support shared decision-making, though it is not precise enough on its own to drive an irreversible clinical decision. It is prognostic, not therapeutic; treatment-effect estimation is addressed in companion work.

Article activity feed