Cross-database validation reveals distinct layers of transportability in ICU delirium prediction

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

External validation of clinical AI emphasizes discrimination, although deployment requires the endpoint, probability estimates and operating policy to transport. Here we show that these layers diverged in retrospective bidirectional evaluation of five model families across eICU and MIMIC-IV. Coarse-label AUROC fell from 0.87–0.92 internally to 0.66–0.83 during source-only transfer. For assessment-conditioned repeated monitoring of persistence or recurrence, external AUROC reached 0.76–0.94, but removing assessment history reduced it by 0.16–0.32; broader features did not help consistently. Transported scores concentrated future-positive ICU stays 2.4–6.9-fold in the top risk decile. Development-selected cutoffs alerted 0.3–2.0% of prediction rows and captured 9.2–11.0% of future-positive rows; after deduplication, 4.9–12.2% of stays were alerted, capturing 43.9–49.4% of future-positive stays. Thus, ranking can persist while probability and policy transport remain site dependent. Layered validation is a prerequisite for prospective evaluation, not evidence of clinical benefit.

Article activity feed