Revision Behavior and Explainability in Adaptive LLM Swarms for ICU Mortality Risk Prediction: A Two-Dataset Evaluation

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Purpose

To evaluate how adaptive LLM swarm revision changes ICU mortality-risk outputs and to characterize evidence use, explanation indicators, auditability, and computational burden in final adaptive outputs relative to an independently executed fixed-voting (FV) architecture.

Methods

We retrospectively analyzed 1,607 eICU encounters (1,500 stays) and 1,607 ICU-2012 encounters. Initial (ASI) and final (ASF) adaptive outputs were compared within runs for revision engagement and risk-score drift; final ASF and separately generated FV outputs were compared for evidence use, explanation indicators, auditability, and computation. Paired differences and 95% confidence intervals used 10,000 hospital-stay-clustered bootstrap replicates.

Results

At least one specialist revision trace occurred in 84.32% of eICU and 56.44% of ICU-2012 encounters. Mean ASF-minus-ASI risk-score changes were +0.0810 and +0.0539 ; 757 of 761 0.50-threshold crossings moved toward mortality, without clear AUROC or AUPRC improvement. In the independent benchmark, ASF explanations contained 1.26 and 0.55 more supporting-evidence items than FV, but counterevidence acknowledgement was 21.59 and 7.47 percentage points lower and unsupported-claim flags were 1.43 and 0.68 points higher. All final records met the reconstruction-completeness criterion, although ASF generated more warnings and required 1.97 and 1.52 times the FV runtime.

Conclusion

Adaptive revision materially changed swarm operating behaviour. Independently, final ASF outputs showed greater supporting-evidence use but less balanced evidence engagement, more process warnings, and greater computational burden than FV. These automated artifact-level findings do not establish superior explanation quality or isolate revision as their cause.

Article activity feed